From 168c37fa5f2ba858ed4ee4de21ce6a54a9170844 Mon Sep 17 00:00:00 2001 From: Peter Kieltyka Date: Fri, 25 Sep 2026 11:37:05 -0400 Subject: [PATCH 1/3] feat(action): named models per comment trigger and provider credential docs (plan 125) Adds a `models` Action input: a flat `alias: provider/model[:reasoning]` list. `model` stays the default (an alias or a full spec) for automatic reviews and a bare trigger; `codegenie review ` runs one review with a listed model. The alias is a closed-set lookup resolved only after the live permission check; an unlisted name gets a fixed reply and no review. Without `models`, behavior is unchanged. `llm-api-key`, when set, is now the only model key: it clears competing credential env vars for its provider and overrides native vars, all configured models must share one provider, and the Action refuses to run when a stored login on a self-hosted runner would override it. Model resolution failures now say unknown model, deprecated model, or missing credentials (naming the env var) instead of one generic message; the Action shows these in the failure comment, which is now sanitized. Syncs the provider env-var table with pi-ai 0.87.1 (qwen-token-plan*, baseten, meta, radius), adds a registry coverage test, generates a Credentials table in models.md (also catching up the stale registry listing), merges the example workflows into one, and updates the README. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01NxBabX2aBd6eRLJryVci78 --- README.md | 52 +- action.yml | 9 +- .../workflows/codegenie-review-comment.yml | 43 - examples/workflows/codegenie-review-pr.yml | 43 - examples/workflows/codegenie-review.yml | 71 + models.md | 1238 +++++++++++------ scripts/write-models-md.mjs | 19 +- ...-model-aliases-and-provider-credentials.md | 271 ++++ specs/plans/README.md | 1 + src/github-action/entrypoint.ts | 201 +-- src/github-action/event-gate.ts | 23 +- src/github-action/models.ts | 203 +++ src/github-action/status-comment.ts | 5 +- src/llm/llm-runner.ts | 12 + src/llm/pi-runner.ts | 71 +- src/provider/pi-ai-models.ts | 49 + tests/github-action.test.ts | 444 +++++- tests/model-resolution.test.ts | 187 +++ 18 files changed, 2330 insertions(+), 612 deletions(-) delete mode 100644 examples/workflows/codegenie-review-comment.yml delete mode 100644 examples/workflows/codegenie-review-pr.yml create mode 100644 examples/workflows/codegenie-review.yml create mode 100644 specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md create mode 100644 src/github-action/models.ts create mode 100644 tests/model-resolution.test.ts diff --git a/README.md b/README.md index 1b016f9..d272ece 100644 --- a/README.md +++ b/README.md @@ -24,6 +24,7 @@ codegenie provider login openai --api-key # OpenAI API key codegenie provider login openrouter --api-key # OpenRouter API key setup # 2. Pick your default model (fuzzy-matched) +codegenie provider use luna:xhigh # -> openai/gpt-6-luna codegenie provider use opus # -> anthropic/claude-opus-5 codegenie provider use gpt-5.5 # -> openai-codex/gpt-5.5 codegenie provider use deepseek-v4.1-flash:max # -> openrouter/deepseek/deepseek-v4.1-flash @@ -107,7 +108,8 @@ jobs: - uses: 0xPolygon/codegenie@v0.6.3 with: # Works with any model! - model: "openrouter/deepseek/deepseek-v4.1-flash:max" + model: "openrouter/openai/gpt-6-luna:xhigh" + # model: "openrouter/deepseek/deepseek-v4.1-flash:max" # model: "openrouter/z-ai/glm-5.3:max" # model: "anthropic/claude-opus-5:high" @@ -117,9 +119,53 @@ jobs: llm-api-key: ${{ secrets.LLM_API_KEY }} ``` -The `model` input is one spec: `provider/model[:reasoning]` — any model in [models.md](./models.md) works (`openai/gpt-5.5:xhigh`, `google/gemini-3-pro`, ...), with reasoning defaulting to `high`. `llm-api-key` is provider-generic: codegenie routes it to whatever variable the named provider reads. Provider-native env vars (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, ...) also work and take precedence if you already keep secrets under those names. +The `model` input is one spec: `provider/model[:reasoning]` — any model in [models.md](./models.md) works (`openai/gpt-5.5:xhigh`, `google/gemini-3-pro`, ...), with reasoning defaulting to `high`. `llm-api-key` is provider-generic: codegenie routes it to whatever variable the named provider reads. When `llm-api-key` is set it is the only model key — it overrides that provider's env vars. Without it, codegenie reads the provider's own env vars (see [Credentials](#credentials)). -See `examples/workflows/` for both trigger lanes (automatic and comment-triggered). All authorization — exact trigger-phrase match, live write-permission check — happens inside codegenie; the workflows contain no gating logic to drift. Cancellation policy is one rule: `cancel-in-progress: true`, newest event wins — a push supersedes the now-stale review. On the comment lane that also means any comment on a PR supersedes that PR's in-flight run before codegenie decides it's a skip; if your PR threads are chatty, set it to `false` there (re-triggers queue instead), or gate a separate ungrouped job with the `preflight-only` input for the strictest setup. Fork `pull_request` events skip cleanly (the comment lane serves fork PRs), and all posting is deterministic harness code — reviewed content and comment text never reach the model as instructions or tools. Costs are the usual two: GitHub Actions minutes and provider tokens. +### Several models, picked per comment + +List named models once. `model` stays the default for automatic reviews and a bare `codegenie review`. A collaborator can comment `codegenie review opus` to run that one review with the `opus` entry instead. + +```yaml + - uses: 0xPolygon/codegenie@v0.6.3 + env: + # one credential env var per provider in the list (see Credentials below) + OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }} + ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} + with: + model: luna # default: an alias below, or a full spec + models: | + luna: openrouter/openai/gpt-6-luna:xhigh + deepseek: openrouter/deepseek/deepseek-v4.1-flash:max + glm: openrouter/z-ai/glm-5.3:max + opus: anthropic/claude-opus-5:high +``` + +- **Grammar.** The first word after the trigger phrase, on the same line, is looked up in `models` (case-insensitive). Nothing else in the comment is read. Comments cannot supply a model spec, a reasoning level or any other option; to offer a reasoning variant, add an alias for it (`opus-max: anthropic/claude-opus-5:max`). Without `models`, text after the trigger phrase is ignored exactly as before. +- **Unknown names.** `codegenie review opsu` runs no review. codegenie replies with the configured names (`Unknown model. Available: luna (default), deepseek, …`), and only to collaborators who pass the permission check. +- **Keys.** `models` holds model specs only — never put keys in it. With several providers, set each provider's env var on the step, as above. If every configured model uses one provider, a single `llm-api-key` covers them all. When `llm-api-key` is set and the models span several providers, every run fails with a configuration error instead of sending one provider's key to another. On a self-hosted runner where someone ran `codegenie provider login` for that provider, the stored login would take precedence, so the run fails and says to log out there or drop `llm-api-key`. +- **Precedence change.** `llm-api-key` now wins over a provider env var for the same provider. Older releases preferred the env var. A workflow that sets both, with different values, now uses `llm-api-key`. +- **One status comment.** A `codegenie review opus` comment supersedes an in-flight run on the same PR (newest event wins), and the PR's single status comment shows the newest report. +- **A mistyped name still cancels.** `codegenie review opsu` starts a run that supersedes the running review, then only posts the reply — the same "any comment supersedes" trade-off described below. +- **Inline comments accumulate across models.** A `codegenie review opus` run after an automatic review posts its own inline findings. Duplicate suppression only matches unchanged wording, so two models reporting the same issue both appear. That is expected for a second opinion. + +### Credentials + +Each provider reads its key from its own env var. The common ones: + +| Provider | Env var | +| --- | --- | +| `anthropic` | `ANTHROPIC_API_KEY` | +| `openai` | `OPENAI_API_KEY` | +| `openrouter` | `OPENROUTER_API_KEY` | +| `google` | **`GEMINI_API_KEY`** | +| `amazon-bedrock` | AWS credentials (`AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY`, `AWS_BEARER_TOKEN_BEDROCK`, or OIDC via `aws-actions/configure-aws-credentials`) | +| `google-vertex` | `GOOGLE_CLOUD_API_KEY`, or Application Default Credentials plus `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION` | + +Most names are the provider id in capitals plus `_API_KEY`. Exceptions include `google` → `GEMINI_API_KEY`, `vercel-ai-gateway` → `AI_GATEWAY_API_KEY`, `huggingface` → `HF_TOKEN`, `github-copilot` → `COPILOT_GITHUB_TOKEN`, and shared names such as `MOONSHOT_API_KEY` and `CLOUDFLARE_API_KEY`. `openai-codex` uses a stored ChatGPT-plan login, so in CI it needs a self-hosted runner with that login. The full table for every provider is in [models.md#credentials](./models.md#credentials). A missing key fails the review with a message naming the env var to set. + +### Triggers, trust and cancellation + +See [`examples/workflows/codegenie-review.yml`](./examples/workflows/codegenie-review.yml) for one workflow that serves both trigger lanes (automatic and comment-triggered). All authorization — exact trigger-phrase match, live write-permission check, model-name lookup — happens inside codegenie; the workflow contains no gating logic to drift. Cancellation policy is one rule: `cancel-in-progress: true`, newest event wins — a push supersedes the now-stale review. On the comment lane that also means any comment on a PR supersedes that PR's in-flight run before codegenie decides it's a skip; if your PR threads are chatty, set it to `false` (re-triggers queue instead), or gate a separate ungrouped job with the `preflight-only` input for the strictest setup. A preflight job needs no model credentials; give it the same `model`, `models` and trigger inputs as the review job. Fork `pull_request` events skip cleanly (the comment lane serves fork PRs), and all posting is deterministic harness code — reviewed content and comment text never reach the model as instructions or tools. Costs are the usual two: GitHub Actions minutes and provider tokens. ## Providers and models diff --git a/action.yml b/action.yml index 2eab531..7e301a3 100644 --- a/action.yml +++ b/action.yml @@ -35,10 +35,13 @@ inputs: description: "Comma-separated review lenses to enable." default: "" model: - description: "Model spec: provider/model[:reasoning], e.g. anthropic/claude-opus-5:xhigh. Reasoning is one of low, medium, high, xhigh, max (default high) and must be a level the model advertises." + description: "Default model for automatic reviews and a bare trigger comment: a provider/model[:reasoning] spec (e.g. anthropic/claude-opus-5:xhigh) or the name of an alias from `models`. Reasoning is one of low, medium, high, xhigh, max (default high) and must be a level the model advertises." + default: "" + models: + description: "Optional named models, one `alias: provider/model[:reasoning]` per line. A collaborator comment ` ` reviews with that model; an unlisted alias gets a reply listing the configured ones. Model specs only — never put keys here." default: "" llm-api-key: - description: "Generic LLM API key, routed to the provider named in `model`. Provider-native env vars (ANTHROPIC_API_KEY, OPENAI_API_KEY, ...) also work and take precedence." + description: "Generic LLM API key. When set it is the only model key: it overrides the provider's env vars, and `model` plus every alias in `models` must use that one provider. When unset, provider env vars (see models.md#credentials) are used." default: "" max-time: description: "Override review.maxTime in minutes for this run." @@ -86,6 +89,7 @@ runs: INPUT_DEPTH: ${{ inputs.depth }} INPUT_LENSES: ${{ inputs.lenses }} INPUT_MODEL: ${{ inputs.model }} + INPUT_MODELS: ${{ inputs.models }} INPUT_BOT_LOGIN: ${{ inputs.bot-login }} INPUT_PREFLIGHT_ONLY: ${{ inputs.preflight-only }} INPUT_MAX_TIME: ${{ inputs.max-time }} @@ -103,6 +107,7 @@ runs: args+=(--post-inline-comments "$INPUT_POST_INLINE_COMMENTS") args+=(--depth "$INPUT_DEPTH") args+=(--model "$INPUT_MODEL") + args+=(--models "$INPUT_MODELS") args+=(--bot-login "$INPUT_BOT_LOGIN") args+=(--preflight-only "$INPUT_PREFLIGHT_ONLY") args+=(--max-time "$INPUT_MAX_TIME") diff --git a/examples/workflows/codegenie-review-comment.yml b/examples/workflows/codegenie-review-comment.yml deleted file mode 100644 index b904bfe..0000000 --- a/examples/workflows/codegenie-review-comment.yml +++ /dev/null @@ -1,43 +0,0 @@ -# Comment-triggered codegenie review: write "codegenie review" on a PR. -# -# issue_comment workflows run from the default branch, so codegenie's config -# and skills come from trusted code and this lane also serves fork PRs. All -# trigger authorization (exact phrase match, live write-permission check) -# happens inside codegenie — no gating logic in YAML. -# -# cancel-in-progress: newest event wins. Note the trade-off: any comment on -# the PR starts a run that supersedes an in-flight review of that PR before -# codegenie decides it is a skip. If your PR threads are chatty, set this to -# false (re-triggers then queue) or gate a separate ungrouped job with the -# `preflight-only` input. -name: codegenie review (comment) - -on: - issue_comment: - types: [created] - -permissions: - contents: read - pull-requests: write - issues: write - -concurrency: - group: codegenie-review-pr-${{ github.event.issue.number }} - cancel-in-progress: true - -jobs: - review: - runs-on: ubuntu-latest - timeout-minutes: 45 - steps: - - uses: actions/checkout@v7 - with: - fetch-depth: 0 - - - uses: 0xPolygon/codegenie@v0.6.3 - with: - model: "openrouter/deepseek/deepseek-v4.1-flash:max" - # model: "openrouter/z-ai/glm-5.3:max" - # model: "anthropic/claude-opus-5:high" - llm-api-key: ${{ secrets.LLM_API_KEY }} - trigger-phrase: "codegenie review" diff --git a/examples/workflows/codegenie-review-pr.yml b/examples/workflows/codegenie-review-pr.yml deleted file mode 100644 index 8eb13e9..0000000 --- a/examples/workflows/codegenie-review-pr.yml +++ /dev/null @@ -1,43 +0,0 @@ -# Automatic codegenie review on every PR open/update. -# -# Uses the `pull_request` event (never `pull_request_target`): fork PRs run -# without secrets and are skipped cleanly — use the comment-trigger workflow -# for fork PRs. Requires the LLM_API_KEY secret (routed to the provider named -# in `model`). -# -# Checkout pins the BASE revision: codegenie's config and skills load from -# trusted code, while the PR head is fetched by `--pr` as untrusted data. -# `cancel-in-progress` is safe here: only pushes to this same PR join the -# group, and a push makes the running review stale by definition. -name: codegenie review - -on: - pull_request: - types: [opened, synchronize, ready_for_review] - -permissions: - contents: read - pull-requests: write - issues: write - -concurrency: - group: codegenie-review-pr-${{ github.event.pull_request.number }} - cancel-in-progress: true - -jobs: - review: - runs-on: ubuntu-latest - timeout-minutes: 45 - steps: - - uses: actions/checkout@v7 - with: - ref: ${{ github.event.pull_request.base.sha }} - fetch-depth: 0 - - - uses: 0xPolygon/codegenie@v0.6.3 - with: - model: "openrouter/deepseek/deepseek-v4.1-flash:max" - # model: "openrouter/z-ai/glm-5.3:max" - # model: "anthropic/claude-opus-5:high" - llm-api-key: ${{ secrets.LLM_API_KEY }} - post-inline-comments: "true" diff --git a/examples/workflows/codegenie-review.yml b/examples/workflows/codegenie-review.yml new file mode 100644 index 0000000..63cc4fc --- /dev/null +++ b/examples/workflows/codegenie-review.yml @@ -0,0 +1,71 @@ +# codegenie review: automatic on every PR open/update, and on demand when a +# collaborator comments "codegenie review" (or "codegenie review " to +# pick one of the models listed below for that run). +# +# Trust model: `pull_request` runs check out the PR's BASE revision, and +# `issue_comment` runs check out the default branch, so codegenie's config and +# skills always come from trusted code; the PR head is fetched by `--pr` as +# untrusted data. Uses `pull_request`, never `pull_request_target`: fork PRs +# run without secrets and skip cleanly — the comment trigger serves them. +# +# All trigger authorization (exact phrase match, live write-permission check, +# model-alias lookup) happens inside codegenie — no gating logic in YAML. +# +# cancel-in-progress: newest event wins. A push makes the running review stale +# by definition. Trade-off: any comment on the PR (including a mistyped model +# name) starts a run that supersedes an in-flight review of that PR before +# codegenie decides it is a skip. If your PR threads are chatty, set this to +# false (re-triggers then queue) or gate a separate ungrouped job with the +# `preflight-only` input. +name: codegenie review + +on: + pull_request: + types: [opened, synchronize, ready_for_review] + issue_comment: + types: [created] + +permissions: + contents: read + pull-requests: write + issues: write + +concurrency: + group: codegenie-review-pr-${{ github.event.pull_request.number || github.event.issue.number }} + cancel-in-progress: true + +jobs: + review: + runs-on: ubuntu-latest + timeout-minutes: 45 + steps: + - uses: actions/checkout@v7 + with: + # base SHA on pull_request events; default branch on comment events + ref: ${{ github.event.pull_request.base.sha || '' }} + fetch-depth: 0 + + - uses: 0xPolygon/codegenie@v0.6.3 + # One credential env var per provider used below; the name for every + # provider is listed in models.md#credentials. + env: + OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }} + ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} + with: + # Default model: automatic reviews and a bare "codegenie review". + model: luna + # "codegenie review " reviews with that model instead. + models: | + luna: openrouter/openai/gpt-6-luna:xhigh + deepseek: openrouter/deepseek/deepseek-v4.1-flash:max + glm: openrouter/z-ai/glm-5.3:max + opus: anthropic/claude-opus-5:high + + # Simple single-model form (no `models`, no `env:` above): + # model: "openrouter/openai/gpt-6-luna:xhigh" + # llm-api-key: ${{ secrets.LLM_API_KEY }} + # When set, llm-api-key is the only model key: it overrides the + # provider env vars, and every configured model must use one provider. + + trigger-phrase: "codegenie review" + post-inline-comments: "true" diff --git a/models.md b/models.md index c02f73e..9235856 100644 --- a/models.md +++ b/models.md @@ -3,7 +3,7 @@ > Generated from the pi model registry ([models.dev](https://models.dev)) — do not edit by hand. > Regenerate with `make models-list`. -codegenie is multi-provider: **1103 models** across **37 providers**. Use any of them as: +codegenie is multi-provider: **1490 models** across **41 providers**. Use any of them as: - `codegenie review --provider --model [--reasoning ]` - `codegenie provider use ` (e.g. `use opus`) to set a default @@ -12,36 +12,97 @@ codegenie is multi-provider: **1103 models** across **37 providers**. Use any of The **Reasoning levels** column shows each model's native thinking levels from the registry (a dash means none). codegenie's `--reasoning` flag (and the `:reasoning` suffix) accepts `low`, `medium`, `high`, `xhigh`, or `auto` and maps onto whatever the model natively supports. Listing here means the model is known, not authenticated — -connect a provider with `codegenie provider login ` (or env vars / the Action's `llm-api-key`). +connect a provider with `codegenie provider login `, or set the env var listed under [Credentials](#credentials). + +## Credentials + +The env vars each provider's credentials are read from, in lookup order. In the GitHub Action, set them on the +`uses:` step's `env:`, or pass one key as `llm-api-key` when every configured model uses the same provider. + +| Provider | Credentials | +| --- | --- | +| `amazon-bedrock` | AWS credentials: `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY`, `AWS_BEARER_TOKEN_BEDROCK`, `AWS_PROFILE`, or an OIDC web identity (e.g. aws-actions/configure-aws-credentials) | +| `ant-ling` | `ANT_LING_API_KEY` | +| `anthropic` | `ANTHROPIC_OAUTH_TOKEN`, `ANTHROPIC_API_KEY` | +| `azure-openai-responses` | `AZURE_OPENAI_API_KEY` | +| `baseten` | `BASETEN_API_KEY` | +| `cerebras` | `CEREBRAS_API_KEY` | +| `cloudflare-ai-gateway` | `CLOUDFLARE_API_KEY` | +| `cloudflare-workers-ai` | `CLOUDFLARE_API_KEY` | +| `deepseek` | `DEEPSEEK_API_KEY` | +| `fireworks` | `FIREWORKS_API_KEY` | +| `github-copilot` | `COPILOT_GITHUB_TOKEN` | +| `google-vertex` | `GOOGLE_CLOUD_API_KEY` or Application Default Credentials plus `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION` (e.g. google-github-actions/auth) | +| `google` | `GEMINI_API_KEY` | +| `groq` | `GROQ_API_KEY` | +| `huggingface` | `HF_TOKEN` | +| `kimi-coding` | `KIMI_API_KEY` | +| `meta` | `META_API_KEY` | +| `minimax-cn` | `MINIMAX_CN_API_KEY` | +| `minimax` | `MINIMAX_API_KEY` | +| `mistral` | `MISTRAL_API_KEY` | +| `moonshotai-cn` | `MOONSHOT_API_KEY` | +| `moonshotai` | `MOONSHOT_API_KEY` | +| `nvidia` | `NVIDIA_API_KEY` | +| `openai-codex` | stored ChatGPT-plan login only (`codegenie provider login openai-codex`); in CI this needs a self-hosted runner with that login | +| `openai` | `OPENAI_API_KEY` | +| `opencode-go` | `OPENCODE_API_KEY` | +| `opencode` | `OPENCODE_API_KEY` | +| `openrouter` | `OPENROUTER_API_KEY` | +| `qwen-token-plan-cn` | `QWEN_TOKEN_PLAN_CN_API_KEY` | +| `qwen-token-plan-individual` | `QWEN_TOKEN_PLAN_API_KEY` | +| `qwen-token-plan` | `QWEN_TOKEN_PLAN_API_KEY` | +| `radius` | `RADIUS_API_KEY` | +| `together` | `TOGETHER_API_KEY` | +| `vercel-ai-gateway` | `AI_GATEWAY_API_KEY` | +| `xai` | `XAI_API_KEY` | +| `xiaomi-token-plan-ams` | `XIAOMI_TOKEN_PLAN_AMS_API_KEY` | +| `xiaomi-token-plan-cn` | `XIAOMI_TOKEN_PLAN_CN_API_KEY` | +| `xiaomi-token-plan-sgp` | `XIAOMI_TOKEN_PLAN_SGP_API_KEY` | +| `xiaomi` | `XIAOMI_API_KEY` | +| `zai-coding-cn` | `ZAI_CODING_CN_API_KEY` | +| `zai` | `ZAI_API_KEY` | ## amazon-bedrock | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `amazon.nova-2-lite-v1:0` | Nova 2 Lite | 128k | 4096 | minimal, low, medium, high | -| `amazon.nova-lite-v1:0` | Nova Lite | 300k | 8192 | — | -| `amazon.nova-micro-v1:0` | Nova Micro | 128k | 8192 | — | -| `amazon.nova-pro-v1:0` | Nova Pro | 300k | 8192 | — | +| `amazon.nova-2-lite-v1:0` | Nova 2 Lite | 1000k | 65535 | minimal, low, medium, high | +| `amazon.nova-lite-v1:0` | Nova Lite | 300k | 10k | — | +| `amazon.nova-micro-v1:0` | Nova Micro | 128k | 10k | — | +| `amazon.nova-pro-v1:0` | Nova Pro | 300k | 10k | — | | `anthropic.claude-fable-5` | Claude Fable 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `anthropic.claude-fable-5-1` | Claude Fable 5.1 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `anthropic.claude-haiku-4-5-20251001-v1:0` | Claude Haiku 4.5 | 200k | 64k | minimal, low, medium, high | | `anthropic.claude-opus-4-1-20250805-v1:0` | Claude Opus 4.1 | 200k | 32k | minimal, low, medium, high | | `anthropic.claude-opus-4-5-20251101-v1:0` | Claude Opus 4.5 | 200k | 64k | minimal, low, medium, high | | `anthropic.claude-opus-4-6-v1` | Claude Opus 4.6 | 1000k | 128k | minimal, low, medium, high, max | | `anthropic.claude-opus-4-7` | Claude Opus 4.7 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `anthropic.claude-opus-4-8` | Claude Opus 4.8 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `anthropic.claude-opus-5-5` | Claude Opus 5.5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `anthropic.claude-sonnet-4-5-20250929-v1:0` | Claude Sonnet 4.5 | 200k | 64k | minimal, low, medium, high | -| `anthropic.claude-sonnet-4-6` | Claude Sonnet 4.6 | 1000k | 64k | minimal, low, medium, high, max | +| `anthropic.claude-sonnet-4-6` | Claude Sonnet 4.6 | 1000k | 128k | minimal, low, medium, high, max | | `anthropic.claude-sonnet-5` | Claude Sonnet 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `apac.amazon.nova-lite-v1:0` | Nova Lite (APAC) | 300k | 10k | — | +| `apac.amazon.nova-micro-v1:0` | Nova Micro (APAC) | 128k | 10k | — | +| `apac.amazon.nova-pro-v1:0` | Nova Pro (APAC) | 300k | 10k | — | +| `apac.anthropic.claude-sonnet-4-20250514-v1:0` | Claude Sonnet 4 (APAC) | 200k | 64k | minimal, low, medium, high | | `au.anthropic.claude-haiku-4-5-20251001-v1:0` | Claude Haiku 4.5 (AU) | 200k | 64k | minimal, low, medium, high | | `au.anthropic.claude-opus-4-6-v1` | AU Anthropic Claude Opus 4.6 | 1000k | 128k | minimal, low, medium, high, max | +| `au.anthropic.claude-opus-4-7` | Claude Opus 4.7 (AU) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `au.anthropic.claude-opus-4-8` | Claude Opus 4.8 (AU) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `au.anthropic.claude-opus-5` | Claude Opus 5 (AU) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `au.anthropic.claude-opus-5-5` | Claude Opus 5.5 (AU) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `au.anthropic.claude-sonnet-4-5-20250929-v1:0` | Claude Sonnet 4.5 (AU) | 200k | 64k | minimal, low, medium, high | | `au.anthropic.claude-sonnet-4-6` | AU Anthropic Claude Sonnet 4.6 | 1000k | 128k | minimal, low, medium, high, max | | `au.anthropic.claude-sonnet-5` | Claude Sonnet 5 (AU) | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `deepseek.r1-v1:0` | DeepSeek-R1 | 128k | 32768 | minimal, low, medium, high | +| `ca.amazon.nova-lite-v1:0` | Nova Lite (CA) | 300k | 10k | — | | `deepseek.v3-v1:0` | DeepSeek-V3.1 | 163840 | 81920 | minimal, low, medium, high | -| `deepseek.v3.2` | DeepSeek-V3.2 | 163840 | 81920 | minimal, low, medium, high | +| `deepseek.v3.2` | DeepSeek V3.2 | 163840 | 81920 | minimal, low, medium, high | +| `eu.amazon.nova-2-lite-v1:0` | Nova 2 Lite (EU) | 1000k | 65535 | minimal, low, medium, high | +| `eu.amazon.nova-lite-v1:0` | Nova Lite (EU) | 300k | 10k | — | +| `eu.amazon.nova-micro-v1:0` | Nova Micro (EU) | 128k | 10k | — | +| `eu.amazon.nova-pro-v1:0` | Nova Pro (EU) | 300k | 10k | — | | `eu.anthropic.claude-fable-5` | Claude Fable 5 (EU) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `eu.anthropic.claude-haiku-4-5-20251001-v1:0` | Claude Haiku 4.5 (EU) | 200k | 64k | minimal, low, medium, high | | `eu.anthropic.claude-opus-4-5-20251101-v1:0` | Claude Opus 4.5 (EU) | 200k | 64k | minimal, low, medium, high | @@ -49,36 +110,53 @@ connect a provider with `codegenie provider login ` (or env vars / the | `eu.anthropic.claude-opus-4-7` | Claude Opus 4.7 (EU) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `eu.anthropic.claude-opus-4-8` | Claude Opus 4.8 (EU) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `eu.anthropic.claude-opus-5` | Claude Opus 5 (EU) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `eu.anthropic.claude-opus-5-5` | Claude Opus 5.5 (EU) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `eu.anthropic.claude-sonnet-4-20250514-v1:0` | Claude Sonnet 4 (EU) | 200k | 64k | minimal, low, medium, high | | `eu.anthropic.claude-sonnet-4-5-20250929-v1:0` | Claude Sonnet 4.5 (EU) | 200k | 64k | minimal, low, medium, high | -| `eu.anthropic.claude-sonnet-4-6` | Claude Sonnet 4.6 (EU) | 1000k | 64k | minimal, low, medium, high, max | +| `eu.anthropic.claude-sonnet-4-6` | Claude Sonnet 4.6 (EU) | 1000k | 128k | minimal, low, medium, high, max | | `eu.anthropic.claude-sonnet-5` | Claude Sonnet 5 (EU) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `eu.mistral.pixtral-large-2502-v1:0` | Pixtral Large (25.02) (EU) | 128k | 8192 | — | +| `global.amazon.nova-2-lite-v1:0` | Nova 2 Lite (Global) | 1000k | 65535 | minimal, low, medium, high | | `global.anthropic.claude-fable-5` | Claude Fable 5 (Global) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `global.anthropic.claude-fable-5-1` | Claude Fable 5.1 (Global) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `global.anthropic.claude-haiku-4-5-20251001-v1:0` | Claude Haiku 4.5 (Global) | 200k | 64k | minimal, low, medium, high | | `global.anthropic.claude-opus-4-5-20251101-v1:0` | Claude Opus 4.5 (Global) | 200k | 64k | minimal, low, medium, high | | `global.anthropic.claude-opus-4-6-v1` | Claude Opus 4.6 (Global) | 1000k | 128k | minimal, low, medium, high, max | | `global.anthropic.claude-opus-4-7` | Claude Opus 4.7 (Global) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `global.anthropic.claude-opus-4-8` | Claude Opus 4.8 (Global) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `global.anthropic.claude-opus-5` | Claude Opus 5 (Global) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `global.anthropic.claude-opus-5-5` | Claude Opus 5.5 (Global) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `global.anthropic.claude-sonnet-4-20250514-v1:0` | Claude Sonnet 4 (Global) | 200k | 64k | minimal, low, medium, high | | `global.anthropic.claude-sonnet-4-5-20250929-v1:0` | Claude Sonnet 4.5 (Global) | 200k | 64k | minimal, low, medium, high | -| `global.anthropic.claude-sonnet-4-6` | Claude Sonnet 4.6 (Global) | 1000k | 64k | minimal, low, medium, high, max | +| `global.anthropic.claude-sonnet-4-6` | Claude Sonnet 4.6 (Global) | 1000k | 128k | minimal, low, medium, high, max | | `global.anthropic.claude-sonnet-5` | Claude Sonnet 5 (Global) | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `google.gemma-3-27b-it` | Google Gemma 3 27B Instruct | 202752 | 8192 | — | -| `google.gemma-3-4b-it` | Gemma 3 4B IT | 128k | 4096 | — | +| `global.openai.gpt-5.6-luna` | GPT-5.6 Luna (Global) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `global.openai.gpt-5.6-sol` | GPT-5.6 Sol (Global) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `global.openai.gpt-5.6-terra` | GPT-5.6 Terra (Global) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `global.openai.gpt-6-astra` | GPT-6 Astra (Global) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `global.xai.grok-4.6` | Grok 4.6 (Global) | 500k | 500k | minimal, low, medium, high | +| `google.gemma-4-26b-a4b` | Gemma 4 26B A4B IT | 262144 | 32768 | minimal, low, medium, high | +| `google.gemma-4-31b` | Gemma 4 31B IT | 262144 | 32768 | minimal, low, medium, high | +| `google.gemma-4-e2b` | Gemma 4 E2B IT | 131072 | 8192 | minimal, low, medium, high | +| `in.openai.gpt-5.6-luna` | GPT-5.6 Luna (India) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `in.openai.gpt-5.6-terra` | GPT-5.6 Terra (India) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `jp.amazon.nova-2-lite-v1:0` | Nova 2 Lite (JP) | 1000k | 65535 | minimal, low, medium, high | | `jp.anthropic.claude-haiku-4-5-20251001-v1:0` | Claude Haiku 4.5 (JP) | 200k | 64k | minimal, low, medium, high | | `jp.anthropic.claude-opus-4-7` | Claude Opus 4.7 (JP) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `jp.anthropic.claude-opus-4-8` | Claude Opus 4.8 (JP) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `jp.anthropic.claude-opus-5` | Claude Opus 5 (JP) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `jp.anthropic.claude-opus-5-5` | Claude Opus 5.5 (JP) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `jp.anthropic.claude-sonnet-4-5-20250929-v1:0` | Claude Sonnet 4.5 (JP) | 200k | 64k | minimal, low, medium, high | -| `jp.anthropic.claude-sonnet-4-6` | Claude Sonnet 4.6 (JP) | 1000k | 64k | minimal, low, medium, high, max | +| `jp.anthropic.claude-sonnet-4-6` | Claude Sonnet 4.6 (JP) | 1000k | 128k | minimal, low, medium, high, max | | `jp.anthropic.claude-sonnet-5` | Claude Sonnet 5 (JP) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `meta.llama3-1-70b-instruct-v1:0` | Llama 3.1 70B Instruct | 128k | 4096 | — | | `meta.llama3-1-8b-instruct-v1:0` | Llama 3.1 8B Instruct | 128k | 4096 | — | | `meta.llama3-3-70b-instruct-v1:0` | Llama 3.3 70B Instruct | 128k | 4096 | — | -| `meta.llama4-maverick-17b-instruct-v1:0` | Llama 4 Maverick 17B Instruct | 1000k | 16384 | — | -| `meta.llama4-scout-17b-instruct-v1:0` | Llama 4 Scout 17B Instruct | 3500k | 16384 | — | -| `minimax.minimax-m2` | MiniMax M2 | 204608 | 128k | minimal, low, medium, high | -| `minimax.minimax-m2.1` | MiniMax M2.1 | 204800 | 131072 | minimal, low, medium, high | -| `minimax.minimax-m2.5` | MiniMax M2.5 | 196608 | 98304 | minimal, low, medium, high | +| `meta.llama4-maverick-17b-instruct-v1:0` | Llama 4 Maverick 17B Instruct | 1000k | 8192 | — | +| `meta.llama4-scout-17b-instruct-v1:0` | Llama 4 Scout 17B Instruct | 10000k | 8192 | — | +| `minimax.minimax-m2` | MiniMax-M2 | 204608 | 128k | minimal, low, medium, high | +| `minimax.minimax-m2.1` | MiniMax-M2.1 | 204800 | 131072 | minimal, low, medium, high | +| `minimax.minimax-m2.5` | MiniMax-M2.5 | 196608 | 98304 | minimal, low, medium, high | | `mistral.devstral-2-123b` | Devstral 2 123B | 256k | 8192 | — | | `mistral.magistral-small-2509` | Magistral Small 1.2 | 128k | 40k | minimal, low, medium, high | | `mistral.ministral-3-14b-instruct` | Ministral 14B 3.0 | 128k | 4096 | — | @@ -86,33 +164,42 @@ connect a provider with `codegenie provider login ` (or env vars / the | `mistral.ministral-3-8b-instruct` | Ministral 3 8B | 128k | 4096 | — | | `mistral.mistral-large-3-675b-instruct` | Mistral Large 3 | 256k | 8192 | — | | `mistral.pixtral-large-2502-v1:0` | Pixtral Large (25.02) | 128k | 8192 | — | -| `mistral.voxtral-mini-3b-2507` | Voxtral Mini 3B 2507 | 128k | 4096 | — | -| `mistral.voxtral-small-24b-2507` | Voxtral Small 24B 2507 | 32k | 8192 | — | +| `mistral.voxtral-mini-3b-2507` | Voxtral Mini 3B 2507 | 32768 | 4096 | — | +| `mistral.voxtral-small-24b-2507` | Voxtral Small 24B 2507 | 32768 | 8192 | — | | `moonshot.kimi-k2-thinking` | Kimi K2 Thinking | 262143 | 16k | minimal, low, medium, high | -| `moonshotai.kimi-k2.5` | Kimi K2.5 | 262143 | 16k | minimal, low, medium, high | -| `nvidia.nemotron-nano-12b-v2` | NVIDIA Nemotron Nano 12B v2 VL BF16 | 128k | 4096 | — | -| `nvidia.nemotron-nano-3-30b` | NVIDIA Nemotron Nano 3 30B | 128k | 4096 | minimal, low, medium, high | -| `nvidia.nemotron-nano-9b-v2` | NVIDIA Nemotron Nano 9B v2 | 128k | 4096 | — | +| `moonshotai.kimi-k2.5` | Kimi K2.5 | 262143 | 16384 | minimal, low, medium, high | +| `nvidia.nemotron-nano-12b-v2` | NVIDIA Nemotron Nano 12B v2 VL BF16 | 128k | 8192 | — | +| `nvidia.nemotron-nano-3-30b` | NVIDIA Nemotron Nano 3 30B | 262144 | 8192 | minimal, low, medium, high | +| `nvidia.nemotron-nano-9b-v2` | NVIDIA Nemotron Nano 9B v2 | 131072 | 8192 | — | | `nvidia.nemotron-super-3-120b` | NVIDIA Nemotron 3 Super 120B A12B | 262144 | 131072 | minimal, low, medium, high | | `openai.gpt-5.4` | GPT-5.4 | 272k | 128k | minimal, low, medium, high, xhigh | | `openai.gpt-5.5` | GPT-5.5 | 272k | 128k | minimal, low, medium, high, xhigh | -| `openai.gpt-5.6-luna` | GPT-5.6 Luna | 272k | 128k | minimal, low, medium, high, xhigh | -| `openai.gpt-5.6-sol` | GPT-5.6 Sol | 272k | 128k | minimal, low, medium, high, xhigh | -| `openai.gpt-5.6-terra` | GPT-5.6 Terra | 272k | 128k | minimal, low, medium, high, xhigh | +| `openai.gpt-5.6-luna` | GPT-5.6 Luna | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai.gpt-5.6-sol` | GPT-5.6 Sol | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai.gpt-5.6-terra` | GPT-5.6 Terra | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai.gpt-6-astra` | GPT-6 Astra | 1050k | 128k | minimal, low, medium, high, xhigh | | `openai.gpt-oss-120b` | gpt-oss-120b | 128k | 16384 | minimal, low, medium, high | | `openai.gpt-oss-120b-1:0` | gpt-oss-120b | 128k | 16384 | minimal, low, medium, high | | `openai.gpt-oss-20b` | gpt-oss-20b | 128k | 16384 | minimal, low, medium, high | | `openai.gpt-oss-20b-1:0` | gpt-oss-20b | 128k | 16384 | minimal, low, medium, high | -| `openai.gpt-oss-safeguard-120b` | GPT OSS Safeguard 120B | 128k | 16384 | — | -| `openai.gpt-oss-safeguard-20b` | GPT OSS Safeguard 20B | 128k | 16384 | — | -| `qwen.qwen3-235b-a22b-2507-v1:0` | Qwen3 235B A22B 2507 | 262144 | 131072 | — | -| `qwen.qwen3-32b-v1:0` | Qwen3 32B (dense) | 16384 | 16384 | minimal, low, medium, high | -| `qwen.qwen3-coder-30b-a3b-v1:0` | Qwen3 Coder 30B A3B Instruct | 262144 | 131072 | — | -| `qwen.qwen3-coder-480b-a35b-v1:0` | Qwen3 Coder 480B A35B Instruct | 131072 | 65536 | — | -| `qwen.qwen3-coder-next` | Qwen3 Coder Next | 131072 | 65536 | minimal, low, medium, high | -| `qwen.qwen3-next-80b-a3b` | Qwen/Qwen3-Next-80B-A3B-Instruct | 262k | 262k | — | -| `qwen.qwen3-vl-235b-a22b` | Qwen/Qwen3-VL-235B-A22B-Instruct | 262k | 262k | — | +| `openai.gpt-oss-safeguard-120b` | GPT OSS Safeguard 120B | 128k | 16384 | minimal, low, medium, high | +| `openai.gpt-oss-safeguard-20b` | GPT OSS Safeguard 20B | 128k | 16384 | minimal, low, medium, high | +| `qwen.qwen3-235b-a22b-2507-v1:0` | Qwen3 235B-A22B Instruct 2507 | 262144 | 131072 | — | +| `qwen.qwen3-32b-v1:0` | Qwen3 32B | 32768 | 16384 | minimal, low, medium, high | +| `qwen.qwen3-coder-30b-a3b-v1:0` | Qwen3-Coder 30B-A3B Instruct | 262144 | 131072 | — | +| `qwen.qwen3-coder-480b-a35b-v1:0` | Qwen3-Coder 480B-A35B Instruct | 131072 | 65536 | — | +| `qwen.qwen3-coder-next` | Qwen3 Coder Next | 262144 | 65536 | — | +| `qwen.qwen3-next-80b-a3b` | Qwen3-Next 80B-A3B Instruct | 262144 | 262k | — | +| `qwen.qwen3-vl-235b-a22b` | Qwen3 VL 235B A22B Instruct | 262144 | 262k | — | +| `us-gov.openai.gpt-oss-120b-1:0` | gpt-oss-120b (GovCloud) | 128k | 16384 | minimal, low, medium, high | +| `us-gov.openai.gpt-oss-20b-1:0` | gpt-oss-20b (GovCloud) | 128k | 16384 | minimal, low, medium, high | +| `us.amazon.nova-2-lite-v1:0` | Nova 2 Lite (US) | 1000k | 65535 | minimal, low, medium, high | +| `us.amazon.nova-lite-v1:0` | Nova Lite (US) | 300k | 10k | — | +| `us.amazon.nova-micro-v1:0` | Nova Micro (US) | 128k | 10k | — | +| `us.amazon.nova-premier-v1:0` | Nova Premier (US) | 1000k | 10k | — | +| `us.amazon.nova-pro-v1:0` | Nova Pro (US) | 300k | 10k | — | | `us.anthropic.claude-fable-5` | Claude Fable 5 (US) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `us.anthropic.claude-fable-5-1` | Claude Fable 5.1 (US) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `us.anthropic.claude-haiku-4-5-20251001-v1:0` | Claude Haiku 4.5 (US) | 200k | 64k | minimal, low, medium, high | | `us.anthropic.claude-opus-4-1-20250805-v1:0` | Claude Opus 4.1 (US) | 200k | 32k | minimal, low, medium, high | | `us.anthropic.claude-opus-4-5-20251101-v1:0` | Claude Opus 4.5 (US) | 200k | 64k | minimal, low, medium, high | @@ -120,18 +207,31 @@ connect a provider with `codegenie provider login ` (or env vars / the | `us.anthropic.claude-opus-4-7` | Claude Opus 4.7 (US) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `us.anthropic.claude-opus-4-8` | Claude Opus 4.8 (US) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `us.anthropic.claude-opus-5` | Claude Opus 5 (US) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `us.anthropic.claude-opus-5-5` | Claude Opus 5.5 (US) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `us.anthropic.claude-sonnet-4-20250514-v1:0` | Claude Sonnet 4 (US) | 200k | 64k | minimal, low, medium, high | | `us.anthropic.claude-sonnet-4-5-20250929-v1:0` | Claude Sonnet 4.5 (US) | 200k | 64k | minimal, low, medium, high | -| `us.anthropic.claude-sonnet-4-6` | Claude Sonnet 4.6 (US) | 1000k | 64k | minimal, low, medium, high, max | +| `us.anthropic.claude-sonnet-4-6` | Claude Sonnet 4.6 (US) | 1000k | 128k | minimal, low, medium, high, max | | `us.anthropic.claude-sonnet-5` | Claude Sonnet 5 (US) | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `us.deepseek.r1-v1:0` | DeepSeek-R1 (US) | 128k | 32768 | minimal, low, medium, high | -| `us.meta.llama4-maverick-17b-instruct-v1:0` | Llama 4 Maverick 17B Instruct (US) | 1000k | 16384 | — | -| `us.meta.llama4-scout-17b-instruct-v1:0` | Llama 4 Scout 17B Instruct (US) | 3500k | 16384 | — | +| `us.meta.llama3-1-70b-instruct-v1:0` | Llama 3.1 70B Instruct (US) | 128k | 4096 | — | +| `us.meta.llama3-1-8b-instruct-v1:0` | Llama 3.1 8B Instruct (US) | 128k | 4096 | — | +| `us.meta.llama3-3-70b-instruct-v1:0` | Llama 3.3 70B Instruct (US) | 128k | 4096 | — | +| `us.meta.llama4-maverick-17b-instruct-v1:0` | Llama 4 Maverick 17B Instruct (US) | 1000k | 8192 | — | +| `us.meta.llama4-scout-17b-instruct-v1:0` | Llama 4 Scout 17B Instruct (US) | 10000k | 8192 | — | +| `us.mistral.pixtral-large-2502-v1:0` | Pixtral Large (25.02) (US) | 128k | 8192 | — | +| `us.openai.gpt-5.6-luna` | GPT-5.6 Luna (US) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `us.openai.gpt-5.6-sol` | GPT-5.6 Sol (US) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `us.openai.gpt-5.6-terra` | GPT-5.6 Terra (US) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `us.openai.gpt-6-astra` | GPT-6 Astra (US) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `us.writer.palmyra-x4-v1:0` | Palmyra X4 (US) | 122880 | 8192 | minimal, low, medium, high | +| `us.writer.palmyra-x5-v1:0` | Palmyra X5 (US) | 1040k | 8192 | minimal, low, medium, high | +| `us.xai.grok-4.6` | Grok 4.6 (US) | 500k | 500k | minimal, low, medium, high | | `writer.palmyra-x4-v1:0` | Palmyra X4 | 122880 | 8192 | minimal, low, medium, high | | `writer.palmyra-x5-v1:0` | Palmyra X5 | 1040k | 8192 | minimal, low, medium, high | | `xai.grok-4.3` | Grok 4.3 | 1000k | 131072 | minimal, low, medium, high | +| `xai.grok-4.6` | Grok 4.6 | 500k | 500k | minimal, low, medium, high | | `zai.glm-4.7` | GLM-4.7 | 204800 | 131072 | minimal, low, medium, high | | `zai.glm-4.7-flash` | GLM-4.7-Flash | 200k | 131072 | minimal, low, medium, high | -| `zai.glm-5` | GLM-5 | 202752 | 101376 | minimal, low, medium, high | +| `zai.glm-5` | GLM-5 | 202752 | 131072 | minimal, low, medium, high | ## ant-ling @@ -146,16 +246,16 @@ connect a provider with `codegenie provider login ` (or env vars / the | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | | `claude-fable-5` | Claude Fable 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-fable-5-1` | Claude Fable 5.1 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `claude-haiku-4-5` | Claude Haiku 4.5 (latest) | 200k | 64k | minimal, low, medium, high | | `claude-haiku-4-5-20251001` | Claude Haiku 4.5 | 200k | 64k | minimal, low, medium, high | -| `claude-opus-4-1` | Claude Opus 4.1 (latest) | 200k | 32k | minimal, low, medium, high | -| `claude-opus-4-1-20250805` | Claude Opus 4.1 | 200k | 32k | minimal, low, medium, high | | `claude-opus-4-5` | Claude Opus 4.5 (latest) | 200k | 64k | minimal, low, medium, high | | `claude-opus-4-5-20251101` | Claude Opus 4.5 | 200k | 64k | minimal, low, medium, high | | `claude-opus-4-6` | Claude Opus 4.6 | 1000k | 128k | minimal, low, medium, high, max | | `claude-opus-4-7` | Claude Opus 4.7 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `claude-opus-4-8` | Claude Opus 4.8 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `claude-opus-5` | Claude Opus 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-opus-5-5` | Claude Opus 5.5 | 1000k | 128k | low, medium, high, xhigh, max | | `claude-sonnet-4-5` | Claude Sonnet 4.5 (latest) | 1000k | 64k | minimal, low, medium, high | | `claude-sonnet-4-5-20250929` | Claude Sonnet 4.5 | 1000k | 64k | minimal, low, medium, high | | `claude-sonnet-4-6` | Claude Sonnet 4.6 | 1000k | 128k | minimal, low, medium, high, max | @@ -196,6 +296,9 @@ connect a provider with `codegenie provider login ` (or env vars / the | `gpt-5.6-luna` | GPT-5.6 Luna | 1050k | 128k | minimal, low, medium, high, xhigh, max | | `gpt-5.6-sol` | GPT-5.6 Sol | 1050k | 128k | minimal, low, medium, high, xhigh, max | | `gpt-5.6-terra` | GPT-5.6 Terra | 1050k | 128k | minimal, low, medium, high, xhigh, max | +| `gpt-6-astra` | GPT-6 Astra | 272k | 128k | low, medium, high, xhigh, max | +| `gpt-6-luna` | GPT-6 Luna | 272k | 128k | low, medium, high, xhigh, max | +| `gpt-6-sol` | GPT-6 Sol | 272k | 128k | low, medium, high, xhigh, max | | `gpt-realtime-2.1` | GPT-Realtime-2.1 | 128k | 32k | minimal, low, medium, high | | `o1` | o1 | 200k | 100k | minimal, low, medium, high | | `o1-pro` | o1-pro | 200k | 100k | minimal, low, medium, high | @@ -204,130 +307,181 @@ connect a provider with `codegenie provider login ` (or env vars / the | `o3-pro` | o3-pro | 200k | 100k | minimal, low, medium, high | | `o4-mini` | o4-mini | 200k | 100k | minimal, low, medium, high | +## baseten + +| Model | Name | Context | Max output | Reasoning levels | +| --- | --- | --- | --- | --- | +| `deepseek-ai/DeepSeek-V4-Flash-0731` | DeepSeek V4 Flash 0731 | 1048576 | 384k | minimal, low, medium, high, xhigh, max | +| `deepseek-ai/DeepSeek-V4-Pro` | DeepSeek V4 Pro | 1048576 | 262144 | minimal, low, medium, high, xhigh, max | +| `deepseek-ai/DeepSeek-V4-Pro-0813` | DeepSeek V4 Pro 0813 | 1048576 | 262144 | low, high, max | +| `deepseek-ai/DeepSeek-V4.1-Flash` | DeepSeek V4.1 Flash | 1048576 | 32768 | low, high, max | +| `moonshotai/Kimi-K2.5` | Kimi K2.5 | 262k | 262k | high | +| `moonshotai/Kimi-K2.6` | Kimi K2.6 | 262k | 262k | high | +| `moonshotai/Kimi-K2.7-Code` | Kimi K2.7 Code | 262k | 262k | high | +| `moonshotai/Kimi-K3` | Kimi K3 | 1048576 | 262144 | low, high, max | +| `nvidia/Nemotron-120B-A12B` | Nemotron Super | 202800 | 202800 | high | +| `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B` | Nemotron Ultra | 202800 | 202800 | high | +| `openai/gpt-oss-120b` | OpenAI GPT 120B | 128072 | 128072 | minimal, low, medium, high, xhigh, max | +| `thinkingmachines/inkling` | Inkling | 1048576 | 32768 | minimal, low, medium, high, xhigh, max | +| `thinkingmachines/inkling-small` | Inkling Small | 1048576 | 32768 | minimal, low, medium, high, xhigh, max | +| `zai-org/GLM-4.7` | GLM 4.7 | 200k | 200k | high | +| `zai-org/GLM-5` | GLM 5 | 202800 | 202800 | high | +| `zai-org/GLM-5.1` | GLM 5.1 | 202800 | 202800 | high | +| `zai-org/GLM-5.2` | GLM 5.2 | 1048576 | 262144 | high, max | +| `zai-org/GLM-5.2-Fast` | GLM 5.2 Fast | 1048576 | 262144 | high, max | +| `zai-org/GLM-5.3` | GLM 5.3 | 1048576 | 262144 | low, high, max | +| `zai-org/GLM-5.3-Fast` | GLM 5.3 Fast | 1048576 | 262144 | low, high, max | +| `zai-org/GLM-5.3-Flash` | GLM 5.3 Flash | 1048576 | 131072 | low, high, max | + ## cerebras | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `gemma-4-31b` | Gemma 4 31B IT | 131072 | 40960 | low, medium, high | | `gpt-oss-120b` | GPT OSS 120B | 131072 | 40960 | low, medium, high | -| `zai-glm-4.7` | Z.AI GLM-4.7 | 131072 | 40960 | — | +| `qwen-3.8-27b` | Qwen3.8 27B | 65536 | 32768 | low, medium, high | ## cloudflare-ai-gateway | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `claude-3-5-haiku` | Claude Haiku 3.5 (latest) | 200k | 8192 | — | -| `claude-3-haiku` | Claude Haiku 3 | 200k | 4096 | — | -| `claude-3-opus` | Claude Opus 3 | 200k | 4096 | — | -| `claude-3-sonnet` | Claude Sonnet 3 | 200k | 4096 | — | -| `claude-3.5-haiku` | Claude Haiku 3.5 (latest) | 200k | 8192 | — | -| `claude-3.5-sonnet` | Claude Sonnet 3.5 v2 | 200k | 8192 | — | | `claude-fable-5` | Claude Fable 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `claude-haiku-4-5` | Claude Haiku 4.5 (latest) | 200k | 64k | minimal, low, medium, high | -| `claude-opus-4` | Claude Opus 4 (latest) | 200k | 32k | minimal, low, medium, high | -| `claude-opus-4-1` | Claude Opus 4.1 (latest) | 200k | 32k | minimal, low, medium, high | -| `claude-opus-4-5` | Claude Opus 4.5 (latest) | 200k | 64k | minimal, low, medium, high | -| `claude-opus-4-6` | Claude Opus 4.6 (latest) | 1000k | 128k | minimal, low, medium, high, max | -| `claude-opus-4-7` | Claude Opus 4.7 | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `claude-opus-4-8` | Claude Opus 4.8 | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `claude-sonnet-4` | Claude Sonnet 4 (latest) | 200k | 64k | minimal, low, medium, high | -| `claude-sonnet-4-5` | Claude Sonnet 4.5 (latest) | 200k | 64k | minimal, low, medium, high | -| `claude-sonnet-4-6` | Claude Sonnet 4.6 | 1000k | 64k | minimal, low, medium, high, max | +| `claude-fable-5.1` | Claude Fable 5.1 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-haiku-4.5` | Claude Haiku 4.5 (latest) | 200k | 64k | minimal, low, medium, high | +| `claude-opus-4.5` | Claude Opus 4.5 (latest) | 200k | 64k | minimal, low, medium, high | +| `claude-opus-4.6` | Claude Opus 4.6 | 1000k | 128k | minimal, low, medium, high, max | +| `claude-opus-4.7` | Claude Opus 4.7 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-opus-4.8` | Claude Opus 4.8 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-opus-5` | Claude Opus 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-sonnet-4.5` | Claude Sonnet 4.5 (latest) | 1000k | 64k | minimal, low, medium, high | +| `claude-sonnet-4.6` | Claude Sonnet 4.6 | 1000k | 128k | minimal, low, medium, high, max | | `claude-sonnet-5` | Claude Sonnet 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `gpt-4` | GPT-4 | 8192 | 8192 | — | -| `gpt-4-turbo` | GPT-4 Turbo | 128k | 4096 | — | +| `gpt-4.1` | GPT-4.1 | 1047576 | 32768 | — | +| `gpt-4.1-mini` | GPT-4.1 mini | 1047576 | 32768 | — | +| `gpt-4.1-nano` | GPT-4.1 nano | 1000k | 32768 | — | | `gpt-4o` | GPT-4o | 128k | 16384 | — | | `gpt-4o-mini` | GPT-4o mini | 128k | 16384 | — | -| `gpt-5.1` | GPT-5.1 | 400k | 128k | low, medium, high | -| `gpt-5.1-codex` | GPT-5.1 Codex | 400k | 128k | low, medium, high | -| `gpt-5.2` | GPT-5.2 | 400k | 128k | low, medium, high, xhigh | -| `gpt-5.2-codex` | GPT-5.2 Codex | 400k | 128k | low, medium, high, xhigh | -| `gpt-5.3-codex` | GPT-5.3 Codex | 400k | 128k | low, medium, high, xhigh | -| `gpt-5.4` | GPT-5.4 | 1050k | 128k | low, medium, high, xhigh | -| `gpt-5.5` | GPT-5.5 | 1050k | 128k | low, medium, high, xhigh | +| `gpt-5` | GPT-5 | 128k | 128k | minimal, low, medium, high | +| `gpt-5-mini` | GPT-5 Mini | 128k | 128k | minimal, low, medium, high | +| `gpt-5-nano` | GPT-5 Nano | 128k | 128k | minimal, low, medium, high | +| `gpt-5.1` | GPT-5.1 | 128k | 128k | low, medium, high | +| `gpt-5.4` | GPT-5.4 | 1000k | 128k | low, medium, high, xhigh | +| `gpt-5.4-mini` | GPT-5.4 mini | 128k | 128k | low, medium, high, xhigh | +| `gpt-5.4-nano` | GPT-5.4 nano | 128k | 128k | low, medium, high, xhigh | +| `gpt-5.4-pro` | GPT-5.4 Pro | 1000k | 128k | medium, high, xhigh | +| `gpt-5.5` | GPT-5.5 | 1000k | 128k | low, medium, high, xhigh | +| `gpt-5.5-pro` | GPT-5.5 Pro | 1000k | 128k | medium, high, xhigh | | `gpt-5.6-luna` | GPT-5.6 Luna | 1050k | 128k | low, medium, high, xhigh, max | | `gpt-5.6-sol` | GPT-5.6 Sol | 1050k | 128k | low, medium, high, xhigh, max | | `gpt-5.6-terra` | GPT-5.6 Terra | 1050k | 128k | low, medium, high, xhigh, max | -| `o1` | o1 | 200k | 100k | low, medium, high | +| `gpt-6-astra` | GPT-6 Astra | 1050k | 128k | low, medium, high, xhigh, max | | `o3` | o3 | 200k | 100k | low, medium, high | | `o3-mini` | o3-mini | 200k | 100k | low, medium, high | -| `o3-pro` | o3-pro | 200k | 100k | low, medium, high | | `o4-mini` | o4-mini | 200k | 100k | low, medium, high | -| `workers-ai/@cf/moonshotai/kimi-k2.5` | Kimi K2.5 | 256k | 256k | minimal, low, medium, high | -| `workers-ai/@cf/moonshotai/kimi-k2.6` | Kimi K2.6 | 256k | 256k | minimal, low, medium, high | +| `workers-ai/@cf/deepseek-ai/deepseek-v4-flash-0731` | DeepSeek V4 Flash 0731 | 1310720 | 1048576 | high, max | +| `workers-ai/@cf/deepseek-ai/deepseek-v4-pro-0813` | DeepSeek V4 Pro 0813 | 1048576 | 1048576 | high, max | +| `workers-ai/@cf/google/gemma-4-26b-a4b-it` | Gemma 4 26B A4B IT | 256k | 16384 | minimal, low, medium, high | +| `workers-ai/@cf/ibm-granite/granite-4.0-h-micro` | Granite 4.0 H Micro | 131k | 131k | — | +| `workers-ai/@cf/meta/llama-3.3-70b-instruct-fp8-fast` | Llama 3.3 70B Instruct fp8 Fast | 24k | 24k | — | +| `workers-ai/@cf/meta/llama-4-scout-17b-16e-instruct` | Llama 4 Scout 17B 16E Instruct | 131k | 16384 | — | +| `workers-ai/@cf/mistralai/mistral-small-3.1-24b-instruct` | Mistral Small 3.1 24B Instruct | 128k | 128k | — | +| `workers-ai/@cf/moonshotai/kimi-k2.6` | Kimi K2.6 | 262144 | 256k | minimal, low, medium, high | +| `workers-ai/@cf/moonshotai/kimi-k2.7-code` | Kimi K2.7 Code | 262144 | 262144 | minimal, low, medium, high | | `workers-ai/@cf/nvidia/nemotron-3-120b-a12b` | Nemotron 3 Super 120B | 256k | 256k | minimal, low, medium, high | +| `workers-ai/@cf/openai/gpt-oss-120b` | GPT OSS 120B | 128k | 16384 | minimal, low, medium, high | +| `workers-ai/@cf/openai/gpt-oss-20b` | GPT OSS 20B | 128k | 16384 | minimal, low, medium, high | +| `workers-ai/@cf/qwen/qwen3-30b-a3b-fp8` | Qwen3 30B A3b fp8 | 32768 | 32768 | minimal, low, medium, high | +| `workers-ai/@cf/qwen/qwen3.8-27b` | Qwen3.8 27B | 262144 | 262144 | minimal, low, medium, high | | `workers-ai/@cf/zai-org/glm-4.7-flash` | GLM-4.7-Flash | 131072 | 131072 | minimal, low, medium, high | -| `workers-ai/@cf/zai-org/glm-5.2` | Glm 5.2 | 262144 | 262144 | minimal, low, medium, high | +| `workers-ai/@cf/zai-org/glm-5.2` | Glm 5.2 | 262144 | 256k | minimal, low, medium, high | +| `workers-ai/@cf/zai-org/glm-5.3` | Glm 5.3 | 1310720 | 1048576 | minimal, low, medium, high | +| `workers-ai/@cf/zai-org/glm-5.3-flash` | Glm 5.3 Flash | 1310720 | 1048576 | minimal, low, medium, high | ## cloudflare-workers-ai | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `@cf/google/gemma-4-26b-a4b-it` | Gemma 4 26B A4B IT | 256k | 16384 | low, medium, high | +| `@cf/deepseek-ai/deepseek-v4-flash-0731` | DeepSeek V4 Flash 0731 | 1310720 | 1048576 | high, max | +| `@cf/deepseek-ai/deepseek-v4-pro-0813` | DeepSeek V4 Pro 0813 | 1048576 | 1048576 | high, max | +| `@cf/google/gemma-4-26b-a4b-it` | Gemma 4 26B A4B IT | 256k | 16384 | high | | `@cf/ibm-granite/granite-4.0-h-micro` | Granite 4.0 H Micro | 131k | 131k | — | | `@cf/meta/llama-3.3-70b-instruct-fp8-fast` | Llama 3.3 70B Instruct fp8 Fast | 24k | 24k | — | | `@cf/meta/llama-4-scout-17b-16e-instruct` | Llama 4 Scout 17B 16E Instruct | 131k | 16384 | — | | `@cf/mistralai/mistral-small-3.1-24b-instruct` | Mistral Small 3.1 24B Instruct | 128k | 128k | — | -| `@cf/moonshotai/kimi-k2.6` | Kimi K2.6 | 262144 | 256k | low, medium, high | -| `@cf/moonshotai/kimi-k2.7-code` | Kimi K2.7 Code | 262144 | 262144 | low, medium, high | +| `@cf/moonshotai/kimi-k2.6` | Kimi K2.6 | 262144 | 256k | high | +| `@cf/moonshotai/kimi-k2.7-code` | Kimi K2.7 Code | 262144 | 262144 | minimal, low, medium, high | | `@cf/nvidia/nemotron-3-120b-a12b` | Nemotron 3 Super 120B | 256k | 256k | low, medium, high | | `@cf/openai/gpt-oss-120b` | GPT OSS 120B | 128k | 16384 | low, medium, high | | `@cf/openai/gpt-oss-20b` | GPT OSS 20B | 128k | 16384 | minimal, low, medium, high | | `@cf/qwen/qwen3-30b-a3b-fp8` | Qwen3 30B A3b fp8 | 32768 | 32768 | minimal, low, medium, high | -| `@cf/zai-org/glm-4.7-flash` | GLM-4.7-Flash | 131072 | 131072 | low, medium, high | -| `@cf/zai-org/glm-5.2` | Glm 5.2 | 262144 | 262144 | low, medium, high | +| `@cf/qwen/qwen3.8-27b` | Qwen3.8 27B | 262144 | 262144 | low, medium, xhigh | +| `@cf/zai-org/glm-4.7-flash` | GLM-4.7-Flash | 131072 | 131072 | minimal, low, medium, high | +| `@cf/zai-org/glm-5.2` | Glm 5.2 | 262144 | 256k | high, max | +| `@cf/zai-org/glm-5.3` | Glm 5.3 | 1310720 | 1048576 | low, high, max | +| `@cf/zai-org/glm-5.3-flash` | Glm 5.3 Flash | 1310720 | 1048576 | low, high, max | ## deepseek | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `deepseek-v4-flash` | DeepSeek V4 Flash | 1000k | 384k | high, max | +| `deepseek-flash` | DeepSeek V4.1 Flash | 1000k | 384k | low, high, max | | `deepseek-v4-pro` | DeepSeek V4 Pro | 1000k | 384k | high, max | ## fireworks | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `accounts/fireworks/models/deepseek-v4-flash` | DeepSeek V4 Flash | 1000k | 384k | minimal, low, medium, high | -| `accounts/fireworks/models/deepseek-v4-pro` | DeepSeek V4 Pro | 1000k | 384k | minimal, low, medium, high | -| `accounts/fireworks/models/glm-5p1` | GLM 5.1 | 202800 | 131072 | minimal, low, medium, high | -| `accounts/fireworks/models/glm-5p2` | GLM 5.2 | 1048575 | 131072 | low, medium, high, max | -| `accounts/fireworks/models/gpt-oss-120b` | GPT OSS 120B | 131072 | 32768 | minimal, low, medium, high | -| `accounts/fireworks/models/gpt-oss-20b` | GPT OSS 20B | 131072 | 32768 | minimal, low, medium, high | +| `accounts/fireworks/models/deepseek-v4-flash-0731` | DeepSeek V4 Flash 0731 | 1000k | 384k | low, high, max | +| `accounts/fireworks/models/deepseek-v4-flash-vision-exp` | DeepSeek V4 Flash Vision Exp | 1000k | 384k | low, high, max | +| `accounts/fireworks/models/deepseek-v4-pro` | DeepSeek V4 Pro | 1000k | 384k | high, max | +| `accounts/fireworks/models/deepseek-v4-pro-0813` | DeepSeek V4 Pro 0813 | 1000k | 384k | low, high, max | +| `accounts/fireworks/models/deepseek-v4p1-flash` | DeepSeek V4.1 Flash | 1000k | 384k | low, high, max | +| `accounts/fireworks/models/glm-5p2` | GLM 5.2 | 1048575 | 131072 | high, max | +| `accounts/fireworks/models/glm-5p3` | GLM 5.3 | 1048573 | 262144 | low, high, max | +| `accounts/fireworks/models/glm-5p3-flash` | GLM 5.3 Flash | 1048573 | 131072 | low, high, max | +| `accounts/fireworks/models/gpt-oss-120b` | GPT OSS 120B | 131072 | 32768 | low, medium, high | +| `accounts/fireworks/models/inkling` | Inkling | 1048576 | 1048576 | minimal, low, medium, high | | `accounts/fireworks/models/kimi-k2p6` | Kimi K2.6 | 262k | 262k | minimal, low, medium, high | | `accounts/fireworks/models/kimi-k2p7-code` | Kimi K2.7 Code | 262k | 262k | minimal, low, medium, high | -| `accounts/fireworks/models/minimax-m2p7` | MiniMax-M2.7 | 196608 | 196608 | minimal, low, medium, high | -| `accounts/fireworks/models/minimax-m3` | MiniMax-M3 | 512k | 512k | minimal, low, medium, high | -| `accounts/fireworks/models/qwen3p7-plus` | Qwen 3.7 Plus | 262144 | 65536 | minimal, low, medium, high | -| `accounts/fireworks/routers/glm-5p1-fast` | GLM 5.1 Fast | 202800 | 131072 | minimal, low, medium, high | -| `accounts/fireworks/routers/glm-5p2-fast` | GLM 5.2 Fast | 1048575 | 131072 | low, medium, high, max | -| `accounts/fireworks/routers/kimi-k2p6-fast` | Kimi K2.6 Fast | 262k | 262k | minimal, low, medium, high | -| `accounts/fireworks/routers/kimi-k2p6-turbo` | Kimi K2.6 Turbo | 262k | 262k | minimal, low, medium, high | -| `accounts/fireworks/routers/kimi-k2p7-code-fast` | Kimi K2.7 Code Fast | 262k | 262k | minimal, low, medium, high | +| `accounts/fireworks/models/kimi-k3` | Kimi K3 | 1048576 | 131072 | low, high, max | +| `accounts/fireworks/models/minimax-m2p7` | MiniMax-M2.7 | 196608 | 131072 | low, medium, high | +| `accounts/fireworks/models/minimax-m3` | MiniMax-M3 | 512k | 512k | low, medium, high | +| `accounts/fireworks/models/muse-glimmer-30b` | Muse Glimmer 30B | 131072 | 131072 | low, medium, high, xhigh | +| `accounts/fireworks/models/nemotron-3-ultra-nvfp4` | Nemotron 3 Ultra 550B A55B | 262144 | 128k | minimal, low, medium, high | +| `accounts/fireworks/models/nemotron-lightning-3p5-30b-a3b` | Nemotron 3.5 Lightning 30B A3B | 262144 | 262144 | minimal, low, medium, high | +| `accounts/fireworks/models/qwen3p7-plus` | Qwen 3.7 Plus | 262144 | 65536 | low, medium, high | +| `accounts/fireworks/models/qwen3p8-2p4t-a95b` | Qwen3.8 2.4T A95B | 262144 | 131072 | low, medium, xhigh | +| `accounts/fireworks/models/qwen3p8-max` | Qwen3.8 Max | 262144 | 131072 | low, medium, xhigh | +| `accounts/fireworks/routers/deepseek-flash-latest` | DeepSeek Flash Latest | 1000k | 384k | low, high, max | +| `accounts/fireworks/routers/deepseek-pro-latest` | DeepSeek Pro Latest | 1000k | 384k | high, max | +| `accounts/fireworks/routers/glm-5p2-fast` | GLM 5.2 Fast | 1048575 | 131072 | high, max | +| `accounts/fireworks/routers/glm-5p3-fast` | GLM 5.3 Fast | 1048572 | 262144 | low, high, max | +| `accounts/fireworks/routers/glm-fast-latest` | GLM 5.3 Fast (Latest) | 1048572 | 262144 | low, high, max | +| `accounts/fireworks/routers/glm-flash-latest` | GLM Flash Latest (GLM 5.3 Flash) | 1048573 | 131072 | low, high, max | +| `accounts/fireworks/routers/glm-latest` | GLM Latest | 1048573 | 262144 | low, high, max | +| `accounts/fireworks/routers/kimi-fast-latest` | Kimi Fast Latest | 1048576 | 131072 | low, medium, high, max | +| `accounts/fireworks/routers/kimi-k3-fast` | Kimi K3 Fast | 1048576 | 131072 | low, high, max | +| `accounts/fireworks/routers/kimi-latest` | Kimi Latest | 1048576 | 131072 | low, medium, high, max | +| `accounts/fireworks/routers/minimax-latest` | MiniMax Latest | 512k | 512k | low, medium, high | +| `accounts/fireworks/routers/qwen-max-latest` | Qwen Max Latest (Qwen3.8 Max) | 262144 | 131072 | minimal, low, medium, high | ## github-copilot | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | | `claude-fable-5` | Claude Fable 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-fable-5.1` | Claude Fable 5.1 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `claude-haiku-4.5` | Claude Haiku 4.5 (latest) | 200k | 64k | minimal, low, medium, high | -| `claude-opus-4.5` | Claude Opus 4.5 (latest) | 200k | 32k | minimal, low, medium, high | -| `claude-opus-4.6` | Claude Opus 4.6 | 1000k | 32k | minimal, low, medium, high, max | | `claude-opus-4.7` | Claude Opus 4.7 | 1000k | 32k | minimal, low, medium, high, xhigh, max | | `claude-opus-4.8` | Claude Opus 4.8 | 1000k | 64k | minimal, low, medium, high, xhigh, max | | `claude-opus-5` | Claude Opus 5 | 1000k | 64k | minimal, low, medium, high, xhigh, max | -| `claude-sonnet-4` | Claude Sonnet 4 (latest) | 216k | 16k | minimal, low, medium, high | -| `claude-sonnet-4.5` | Claude Sonnet 4.5 (latest) | 200k | 32k | minimal, low, medium, high | +| `claude-opus-5.5` | Claude Opus 5.5 | 1000k | 128k | low, medium, high, xhigh, max | | `claude-sonnet-4.6` | Claude Sonnet 4.6 | 1000k | 32k | minimal, low, medium, high, max | | `claude-sonnet-5` | Claude Sonnet 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `gemini-2.5-pro` | Gemini 2.5 Pro | 128k | 64k | minimal, low, medium, high | -| `gemini-3-flash-preview` | Gemini 3 Flash Preview | 128k | 64k | minimal, low, medium, high | -| `gemini-3.1-pro-preview` | Gemini 3.1 Pro Preview | 1000k | 64k | minimal, low, medium, high | | `gemini-3.5-flash` | Gemini 3.5 Flash | 200k | 64k | minimal, low, medium, high | -| `gpt-4.1` | GPT-4.1 | 128k | 16384 | — | +| `gemini-3.6-flash` | Gemini 3.6 Flash | 1000k | 64k | minimal, low, medium, high | +| `gemini-3.7-flash` | Gemini 3.7 Flash | 1000k | 64k | minimal, low, medium, high | +| `gemini-3.8-flash` | Gemini 3.8 Flash | 1000k | 64k | minimal, low, medium, high | | `gpt-5-mini` | GPT-5 Mini | 264k | 64k | minimal, low, medium, high | -| `gpt-5.2` | GPT-5.2 | 400k | 128k | minimal, low, medium, high, xhigh | -| `gpt-5.2-codex` | GPT-5.2 Codex | 400k | 128k | minimal, low, medium, high, xhigh | | `gpt-5.3-codex` | GPT-5.3 Codex | 1000k | 128k | minimal, low, medium, high, xhigh | | `gpt-5.4` | GPT-5.4 | 1000k | 128k | minimal, low, medium, high, xhigh | | `gpt-5.4-mini` | GPT-5.4 mini | 400k | 128k | minimal, low, medium, high, xhigh | @@ -336,8 +490,16 @@ connect a provider with `codegenie provider login ` (or env vars / the | `gpt-5.6-luna` | GPT-5.6 Luna | 1050k | 128k | minimal, low, medium, high, xhigh, max | | `gpt-5.6-sol` | GPT-5.6 Sol | 1050k | 128k | minimal, low, medium, high, xhigh, max | | `gpt-5.6-terra` | GPT-5.6 Terra | 1050k | 128k | minimal, low, medium, high, xhigh, max | +| `gpt-6-astra` | GPT-6 Astra | 1000k | 128k | low, medium, high, xhigh, max | +| `gpt-6-luna` | GPT-6 Luna | 1000k | 128k | low, medium, high, xhigh, max | +| `gpt-6-sol` | GPT-6 Sol | 1000k | 128k | low, medium, high, xhigh, max | +| `grok-4.5` | Grok 4.5 | 500k | 128k | low, medium, high | +| `grok-4.6` | Grok 4.6 | 500k | 128k | low, medium, high, xhigh | +| `grok-4.7` | Grok 4.7 | 500k | 128k | low, medium, high, xhigh | | `kimi-k2.7-code` | Kimi K2.7 Code | 256k | 32k | minimal, low, medium, high | +| `kimi-k3` | Kimi K3 | 1048576 | 131072 | minimal, low, medium, high | | `mai-code-1-flash-picker` | MAI-Code-1-Flash | 256k | 128k | low, medium, high | +| `mai-code-1.1-flash` | MAI-Code-1.1-Flash | 256k | 128k | low, medium, high | ## google-vertex @@ -348,11 +510,13 @@ connect a provider with `codegenie provider login ` (or env vars / the | `gemini-2.5-pro` | Gemini 2.5 Pro | 1048576 | 65536 | minimal, low, medium, high | | `gemini-3-flash-preview` | Gemini 3 Flash Preview | 1048576 | 65536 | minimal, low, medium, high | | `gemini-3.1-flash-lite` | Gemini 3.1 Flash Lite | 1048576 | 65536 | minimal, low, medium, high | -| `gemini-3.1-pro-preview` | Gemini 3.1 Pro Preview | 1048576 | 65536 | low, high | -| `gemini-3.1-pro-preview-customtools` | Gemini 3.1 Pro Preview Custom Tools | 1048576 | 65536 | low, high | +| `gemini-3.1-pro-preview` | Gemini 3.1 Pro Preview | 1048576 | 65536 | low, medium, high | +| `gemini-3.1-pro-preview-customtools` | Gemini 3.1 Pro Preview Custom Tools | 1048576 | 65536 | low, medium, high | | `gemini-3.5-flash` | Gemini 3.5 Flash | 1048576 | 65536 | minimal, low, medium, high | | `gemini-3.5-flash-lite` | Gemini 3.5 Flash Lite | 1048576 | 65536 | minimal, low, medium, high | | `gemini-3.6-flash` | Gemini 3.6 Flash | 1048576 | 65536 | minimal, low, medium, high | +| `gemini-3.7-flash` | Gemini 3.7 Flash | 1048576 | 65536 | low, medium, high | +| `gemini-3.8-flash` | Gemini 3.8 Flash | 1048576 | 65536 | low, medium, high | | `gemini-flash-latest` | Gemini Flash Latest | 1048576 | 65536 | minimal, low, medium, high | | `gemini-flash-lite-latest` | Gemini Flash-Lite Latest | 1048576 | 65536 | minimal, low, medium, high | @@ -362,26 +526,24 @@ connect a provider with `codegenie provider login ` (or env vars / the | --- | --- | --- | --- | --- | | `deep-research-max-preview-04-2026` | Deep Research Max Preview (Apr-21-2026) | 131072 | 65536 | minimal, low, medium, high | | `deep-research-preview-04-2026` | Deep Research Preview (Apr-21-2026) | 131072 | 65536 | minimal, low, medium, high | -| `gemini-2.0-flash` | Gemini 2.0 Flash | 1048576 | 8192 | — | -| `gemini-2.0-flash-lite` | Gemini 2.0 Flash-Lite | 1048576 | 8192 | — | | `gemini-2.5-computer-use-preview-10-2025` | Gemini 2.5 Computer Use Preview 10-2025 | 131072 | 65536 | minimal, low, medium, high | | `gemini-2.5-flash` | Gemini 2.5 Flash | 1048576 | 65536 | minimal, low, medium, high | | `gemini-2.5-flash-lite` | Gemini 2.5 Flash-Lite | 1048576 | 65536 | minimal, low, medium, high | | `gemini-2.5-pro` | Gemini 2.5 Pro | 1048576 | 65536 | minimal, low, medium, high | | `gemini-3-flash-preview` | Gemini 3 Flash Preview | 1048576 | 65536 | minimal, low, medium, high | -| `gemini-3-pro-preview` | Gemini 3 Pro Preview | 1048576 | 65536 | low, high | | `gemini-3.1-flash-lite` | Gemini 3.1 Flash Lite | 1048576 | 65536 | minimal, low, medium, high | -| `gemini-3.1-flash-lite-image` | Nano Banana 2 Lite | 65536 | 65536 | minimal, low, medium, high | +| `gemini-3.1-flash-lite-image` | Nano Banana 2 Lite | 65536 | 65536 | minimal, high | | `gemini-3.1-flash-lite-preview` | Gemini 3.1 Flash Lite Preview | 1048576 | 65536 | minimal, low, medium, high | | `gemini-3.1-flash-live-preview` | Gemini 3.1 Flash Live Preview | 131072 | 65536 | minimal, low, medium, high | -| `gemini-3.1-pro-preview` | Gemini 3.1 Pro Preview | 1048576 | 65536 | low, high | -| `gemini-3.1-pro-preview-customtools` | Gemini 3.1 Pro Preview Custom Tools | 1048576 | 65536 | low, high | +| `gemini-3.1-pro-preview` | Gemini 3.1 Pro Preview | 1048576 | 65536 | low, medium, high | +| `gemini-3.1-pro-preview-customtools` | Gemini 3.1 Pro Preview Custom Tools | 1048576 | 65536 | low, medium, high | | `gemini-3.5-flash` | Gemini 3.5 Flash | 1048576 | 65536 | minimal, low, medium, high | | `gemini-3.5-flash-lite` | Gemini 3.5 Flash Lite | 1048576 | 65536 | minimal, low, medium, high | | `gemini-3.6-flash` | Gemini 3.6 Flash | 1048576 | 65536 | minimal, low, medium, high | +| `gemini-3.7-flash` | Gemini 3.7 Flash | 1048576 | 65536 | low, medium, high | +| `gemini-3.8-flash` | Gemini 3.8 Flash | 1048576 | 65536 | low, medium, high | | `gemini-flash-latest` | Gemini Flash Latest | 1048576 | 65536 | minimal, low, medium, high | | `gemini-flash-lite-latest` | Gemini Flash-Lite Latest | 1048576 | 65536 | minimal, low, medium, high | -| `gemini-robotics-er-1.6-preview` | Gemini Robotics-ER 1.6 Preview | 131072 | 65536 | minimal, low, medium, high | | `gemma-4-26b-a4b-it` | Gemma 4 26B A4B IT | 262144 | 32768 | minimal, high | | `gemma-4-31b-it` | Gemma 4 31B IT | 262144 | 32768 | minimal, high | @@ -391,11 +553,11 @@ connect a provider with `codegenie provider login ` (or env vars / the | --- | --- | --- | --- | --- | | `llama-3.1-8b-instant` | Llama 3.1 8B | 131072 | 131072 | — | | `llama-3.3-70b-versatile` | Llama 3.3 70B | 131072 | 32768 | — | -| `meta-llama/llama-4-scout-17b-16e-instruct` | Llama 4 Scout 17B 16E | 131072 | 8192 | — | | `openai/gpt-oss-120b` | GPT OSS 120B | 131072 | 65536 | low, medium, high | | `openai/gpt-oss-20b` | GPT OSS 20B | 131072 | 65536 | low, medium, high | | `openai/gpt-oss-safeguard-20b` | Safety GPT OSS 20B | 131072 | 65536 | low, medium, high | -| `qwen/qwen3-32b` | Qwen3-32B | 131072 | 40960 | high | +| `qwen/qwen3.6-27b` | Qwen3.6 27B | 131072 | 16384 | high | +| `qwen/qwen3.8-27b` | Qwen3.8 27B | 131042 | 16384 | low, medium, high | ## huggingface @@ -403,33 +565,50 @@ connect a provider with `codegenie provider login ` (or env vars / the | --- | --- | --- | --- | --- | | `deepseek-ai/DeepSeek-R1` | DeepSeek-R1 | 64k | 32768 | minimal, low, medium, high | | `deepseek-ai/DeepSeek-R1-0528` | DeepSeek-R1-0528 | 163840 | 163840 | minimal, low, medium, high | +| `deepseek-ai/DeepSeek-V3` | DeepSeek-V3 | 64k | 8192 | — | +| `deepseek-ai/DeepSeek-V3-0324` | DeepSeek V3 0324 | 163840 | 163840 | — | +| `deepseek-ai/DeepSeek-V3.1` | DeepSeek-V3.1 | 131072 | 8192 | minimal, low, medium, high | | `deepseek-ai/DeepSeek-V3.2` | DeepSeek-V3.2 | 163840 | 65536 | minimal, low, medium, high | | `deepseek-ai/DeepSeek-V4-Flash` | DeepSeek V4 Flash | 1048576 | 384k | minimal, low, medium, high | +| `deepseek-ai/DeepSeek-V4-Flash-0731` | DeepSeek V4 Flash 0731 | 1048576 | 384k | high, max | +| `deepseek-ai/DeepSeek-V4-Flash-Vision-Exp` | DeepSeek V4 Flash Vision Exp | 1048576 | 384k | low, high, max | | `deepseek-ai/DeepSeek-V4-Pro` | DeepSeek V4 Pro | 1048576 | 393216 | high | +| `deepseek-ai/DeepSeek-V4-Pro-0813` | DeepSeek V4 Pro 0813 | 1000k | 384k | low, high, max | +| `deepseek-ai/DeepSeek-V4.1-Flash` | DeepSeek V4.1 Flash | 1048576 | 384k | low, high, xhigh, max | +| `google/gemma-3-12b-it` | Gemma 3 12B IT | 131072 | 131072 | — | +| `google/gemma-3-27b-it` | Gemma 3 27B IT | 131072 | 131072 | — | +| `google/gemma-3-4b-it` | Gemma 3 4B IT | 131072 | 131072 | — | | `google/gemma-4-26B-A4B-it` | Gemma 4 26B A4B IT | 262144 | 32768 | minimal, low, medium, high | | `google/gemma-4-31B-it` | Gemma 4 31B IT | 262144 | 32768 | minimal, low, medium, high | +| `meta-llama/Llama-3.1-8B-Instruct` | Llama-3.1-8B-Instruct | 131072 | 4096 | — | | `meta-llama/Llama-3.3-70B-Instruct` | Llama-3.3-70B-Instruct | 131072 | 4096 | — | -| `MiniMaxAI/MiniMax-M2` | MiniMax-M2 | 204800 | 128k | minimal, low, medium, high | +| `MiniMaxAI/MiniMax-M2` | MiniMax-M2 | 204800 | 131072 | minimal, low, medium, high | | `MiniMaxAI/MiniMax-M2.1` | MiniMax-M2.1 | 204800 | 131072 | minimal, low, medium, high | | `MiniMaxAI/MiniMax-M2.5` | MiniMax-M2.5 | 204800 | 131072 | minimal, low, medium, high | | `MiniMaxAI/MiniMax-M2.7` | MiniMax-M2.7 | 204800 | 131072 | minimal, low, medium, high | -| `MiniMaxAI/MiniMax-M3` | MiniMax-M3 | 524288 | 128k | minimal, low, medium, high | +| `MiniMaxAI/MiniMax-M3` | MiniMax-M3 | 524288 | 512k | minimal, low, medium, high | | `moonshotai/Kimi-K2-Instruct` | Kimi-K2-Instruct | 131072 | 16384 | — | | `moonshotai/Kimi-K2-Instruct-0905` | Kimi-K2-Instruct-0905 | 262144 | 16384 | — | | `moonshotai/Kimi-K2-Thinking` | Kimi-K2-Thinking | 262144 | 262144 | minimal, low, medium, high | | `moonshotai/Kimi-K2.5` | Kimi-K2.5 | 262144 | 262144 | minimal, low, medium, high | | `moonshotai/Kimi-K2.6` | Kimi-K2.6 | 262144 | 262144 | minimal, low, medium, high | | `moonshotai/Kimi-K2.7-Code` | Kimi K2.7 Code | 262144 | 262144 | minimal, low, medium, high | +| `moonshotai/Kimi-K3` | Kimi K3 | 1000k | 131072 | low, high, max | | `openai/gpt-oss-120b` | GPT OSS 120B | 131072 | 32768 | low, medium, high | | `openai/gpt-oss-20b` | GPT OSS 20B | 131072 | 32768 | low, medium, high | +| `Qwen/Qwen2.5-Coder-32B-Instruct` | Qwen2.5-Coder-32B-Instruct | 131072 | 8192 | — | | `Qwen/Qwen3-235B-A22B` | Qwen3 235B-A22B | 40960 | 16384 | minimal, low, medium, high | +| `Qwen/Qwen3-235B-A22B-Instruct-2507` | Qwen3 235B-A22B Instruct 2507 | 262144 | 16384 | — | | `Qwen/Qwen3-235B-A22B-Thinking-2507` | Qwen3-235B-A22B-Thinking-2507 | 262144 | 131072 | minimal, low, medium, high | +| `Qwen/Qwen3-30B-A3B` | Qwen3 30B A3B | 40960 | 16384 | minimal, low, medium, high | | `Qwen/Qwen3-32B` | Qwen3 32B | 131072 | 16384 | minimal, low, medium, high | | `Qwen/Qwen3-Coder-30B-A3B-Instruct` | Qwen3-Coder 30B-A3B Instruct | 262144 | 65536 | — | | `Qwen/Qwen3-Coder-480B-A35B-Instruct` | Qwen3-Coder-480B-A35B-Instruct | 262144 | 66536 | — | | `Qwen/Qwen3-Coder-Next` | Qwen3-Coder-Next | 262144 | 65536 | — | | `Qwen/Qwen3-Next-80B-A3B-Instruct` | Qwen3-Next-80B-A3B-Instruct | 262144 | 66536 | — | | `Qwen/Qwen3-Next-80B-A3B-Thinking` | Qwen3-Next-80B-A3B-Thinking | 262144 | 131072 | — | +| `Qwen/Qwen3-VL-235B-A22B-Instruct` | Qwen3 VL 235B A22B Instruct | 131072 | 32768 | — | +| `Qwen/Qwen3-VL-235B-A22B-Thinking` | Qwen3 VL 235B A22B Thinking | 131072 | 32768 | low, medium, high | | `Qwen/Qwen3.5-122B-A10B` | Qwen3.5 122B-A10B | 262144 | 65536 | minimal, low, medium, high | | `Qwen/Qwen3.5-27B` | Qwen3.5 27B | 262144 | 65536 | minimal, low, medium, high | | `Qwen/Qwen3.5-35B-A3B` | Qwen3.5 35B-A3B | 262144 | 65536 | minimal, low, medium, high | @@ -437,8 +616,14 @@ connect a provider with `codegenie provider login ` (or env vars / the | `Qwen/Qwen3.5-9B` | Qwen3.5 9B | 262144 | 65536 | minimal, low, medium, high | | `Qwen/Qwen3.6-27B` | Qwen3.6 27B | 262144 | 65536 | minimal, low, medium, high | | `Qwen/Qwen3.6-35B-A3B` | Qwen3.6 35B-A3B | 262144 | 65536 | minimal, low, medium, high | +| `Qwen/Qwen3.8-2.4T-A95B` | Qwen3.8 2.4T A95B | 262144 | 131072 | low, medium, xhigh | +| `Qwen/Qwen3.8-27B` | Qwen3.8 27B | 262144 | 32768 | low, medium, xhigh | | `stepfun-ai/Step-3.5-Flash` | Step 3.5 Flash | 262144 | 256k | minimal, low, medium, high | | `stepfun-ai/Step-3.7-Flash` | Step 3.7 Flash | 262144 | 256k | low, medium, high | +| `tencent/Hy3` | Hy3 | 262144 | 128k | low, high | +| `tencent/Hy4-preview` | Hy4 preview | 1000k | 64k | high | +| `thinkingmachines/Inkling` | Inkling | 1048576 | 1048576 | low, medium, high | +| `thinkingmachines/Inkling-Small` | Inkling Small | 524288 | 1048576 | minimal, low, medium, high | | `XiaomiMiMo/MiMo-V2-Flash` | MiMo-V2-Flash | 262144 | 4096 | minimal, low, medium, high | | `XiaomiMiMo/MiMo-V2.5` | MiMo-V2.5 | 262144 | 131072 | low, medium, high, xhigh | | `XiaomiMiMo/MiMo-V2.5-Pro` | MiMo-V2.5-Pro | 1048576 | 131072 | low, medium, high, xhigh | @@ -446,11 +631,14 @@ connect a provider with `codegenie provider login ` (or env vars / the | `zai-org/GLM-4.5-Air` | GLM-4.5-Air | 131072 | 98304 | minimal, low, medium, high | | `zai-org/GLM-4.5V` | GLM-4.5V | 65536 | 16384 | minimal, low, medium, high | | `zai-org/GLM-4.6` | GLM-4.6 | 204800 | 131072 | minimal, low, medium, high | +| `zai-org/GLM-4.6V-Flash` | GLM-4.6V-Flash | 131072 | 32768 | minimal, low, medium, high | | `zai-org/GLM-4.7` | GLM-4.7 | 204800 | 131072 | minimal, low, medium, high | | `zai-org/GLM-4.7-Flash` | GLM-4.7-Flash | 200k | 128k | minimal, low, medium, high | | `zai-org/GLM-5` | GLM-5 | 202752 | 131072 | minimal, low, medium, high | | `zai-org/GLM-5.1` | GLM-5.1 | 202752 | 131072 | minimal, low, medium, high | | `zai-org/GLM-5.2` | GLM-5.2 | 262144 | 131072 | minimal, low, medium, high | +| `zai-org/GLM-5.3` | GLM-5.3 | 1048576 | 131072 | low, high, max | +| `zai-org/GLM-5.3-Flash` | GLM-5.3-Flash | 1048576 | 131072 | low, high, max | ## kimi-coding @@ -458,16 +646,26 @@ connect a provider with `codegenie provider login ` (or env vars / the | --- | --- | --- | --- | --- | | `k3` | Kimi K3 | 1048576 | 131072 | low, high, max | | `k3-256k` | Kimi K3-256K | 262144 | 131072 | low, high, max | -| `kimi-for-coding` | Kimi K2.7 Code | 262144 | 32768 | minimal, low, medium, high | +| `kimi-for-coding` | kimi-for-coding | 1048576 | 32768 | low, high, max | | `kimi-for-coding-highspeed` | Kimi For Coding HighSpeed | 262144 | 32768 | minimal, low, medium, high | +## meta + +| Model | Name | Context | Max output | Reasoning levels | +| --- | --- | --- | --- | --- | +| `muse-spark-1.1` | Muse Spark 1.1 | 1048576 | 131072 | minimal, low, medium, high, xhigh | +| `muse-spark-1.2` | Muse Spark 1.2 | 1048576 | 131072 | minimal, low, medium, high, xhigh | +| `muse-spark-1.2-contributor` | Muse Spark 1.2 Contributor | 1048576 | 131072 | minimal, low, medium, high, xhigh | +| `muse-spark-1.3` | Muse Spark 1.3 | 1048576 | 131072 | minimal, low, medium, high, xhigh, max | +| `muse-spark-1.3-contributor` | Muse Spark 1.3 Contributor | 1048576 | 131072 | minimal, low, medium, high, xhigh | + ## minimax-cn | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | | `MiniMax-M2.7` | MiniMax-M2.7 | 204800 | 131072 | minimal, low, medium, high | | `MiniMax-M2.7-highspeed` | MiniMax-M2.7-highspeed | 204800 | 131072 | minimal, low, medium, high | -| `MiniMax-M3` | MiniMax-M3 | 1000k | 128k | minimal, low, medium, high | +| `MiniMax-M3` | MiniMax-M3 | 1048576 | 512k | minimal, low, medium, high | ## minimax @@ -475,7 +673,7 @@ connect a provider with `codegenie provider login ` (or env vars / the | --- | --- | --- | --- | --- | | `MiniMax-M2.7` | MiniMax-M2.7 | 204800 | 131072 | minimal, low, medium, high | | `MiniMax-M2.7-highspeed` | MiniMax-M2.7-highspeed | 204800 | 131072 | minimal, low, medium, high | -| `MiniMax-M3` | MiniMax-M3 | 1000k | 128k | minimal, low, medium, high | +| `MiniMax-M3` | MiniMax-M3 | 1048576 | 512k | minimal, low, medium, high | ## mistral @@ -511,17 +709,14 @@ connect a provider with `codegenie provider login ` (or env vars / the | `open-mixtral-8x7b` | Mixtral 8x7B | 32k | 32k | — | | `pixtral-12b` | Pixtral 12B | 128k | 128k | — | | `pixtral-large-latest` | Pixtral Large (latest) | 128k | 128k | — | +| `voxtral-small-latest` | Voxtral Small (latest) | 32k | 32k | — | +| `zai-glm-5-2` | GLM-5.2 | 1000k | 131072 | minimal, low, medium, high | +| `zai-glm-5-3` | GLM-5.3 | 1000k | 131072 | minimal, low, medium, high | ## moonshotai-cn | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `kimi-k2-0711-preview` | Kimi K2 0711 | 131072 | 16384 | — | -| `kimi-k2-0905-preview` | Kimi K2 0905 | 262144 | 262144 | — | -| `kimi-k2-thinking` | Kimi K2 Thinking | 262144 | 262144 | minimal, low, medium, high | -| `kimi-k2-thinking-turbo` | Kimi K2 Thinking Turbo | 262144 | 262144 | minimal, low, medium, high | -| `kimi-k2-turbo-preview` | Kimi K2 Turbo | 262144 | 262144 | — | -| `kimi-k2.5` | Kimi K2.5 | 262144 | 262144 | minimal, low, medium, high | | `kimi-k2.6` | Kimi K2.6 | 262144 | 262144 | minimal, low, medium, high | | `kimi-k2.7-code` | Kimi K2.7 Code | 262144 | 262144 | minimal, low, medium, high | | `kimi-k2.7-code-highspeed` | Kimi K2.7 Code HighSpeed | 262144 | 262144 | minimal, low, medium, high | @@ -531,12 +726,6 @@ connect a provider with `codegenie provider login ` (or env vars / the | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `kimi-k2-0711-preview` | Kimi K2 0711 | 131072 | 16384 | — | -| `kimi-k2-0905-preview` | Kimi K2 0905 | 262144 | 262144 | — | -| `kimi-k2-thinking` | Kimi K2 Thinking | 262144 | 262144 | minimal, low, medium, high | -| `kimi-k2-thinking-turbo` | Kimi K2 Thinking Turbo | 262144 | 262144 | minimal, low, medium, high | -| `kimi-k2-turbo-preview` | Kimi K2 Turbo | 262144 | 262144 | — | -| `kimi-k2.5` | Kimi K2.5 | 262144 | 262144 | minimal, low, medium, high | | `kimi-k2.6` | Kimi K2.6 | 262144 | 262144 | minimal, low, medium, high | | `kimi-k2.7-code` | Kimi K2.7 Code | 262144 | 262144 | minimal, low, medium, high | | `kimi-k2.7-code-highspeed` | Kimi K2.7 Code HighSpeed | 262144 | 262144 | minimal, low, medium, high | @@ -546,36 +735,38 @@ connect a provider with `codegenie provider login ` (or env vars / the | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `meta/llama-3.1-70b-instruct` | Llama 3.1 70b Instruct | 128k | 4096 | — | -| `meta/llama-3.1-8b-instruct` | Llama 3.1 8B Instruct | 16k | 4096 | — | +| `google/gemma-3-12b-it` | Gemma 3 12B IT | 131072 | 16384 | — | +| `google/gemma-3-4b-it` | Gemma 3 4B IT | 131072 | 16384 | — | | `meta/llama-3.2-11b-vision-instruct` | Llama 3.2 11b Vision Instruct | 128k | 4096 | — | | `meta/llama-3.2-90b-vision-instruct` | Llama-3.2-90B-Vision-Instruct | 128k | 8192 | — | -| `meta/llama-3.3-70b-instruct` | Llama 3.3 70b Instruct | 128k | 4096 | — | -| `minimaxai/minimax-m3` | MiniMax-M3 | 1000k | 16384 | minimal, low, medium, high | -| `mistralai/mistral-small-4-119b-2603` | mistral-small-4-119b-2603 | 128k | 8192 | minimal, low, medium, high | +| `meta/muse-glimmer-30b` | Muse Glimmer 30B | 131072 | 131072 | minimal, low, medium, high | +| `mistralai/mistral-7b-instruct-v0.3` | Mistral-7B-Instruct-v0.3 | 65536 | 65536 | — | | `moonshotai/kimi-k2.6` | Kimi K2.6 | 262144 | 262144 | minimal, low, medium, high | -| `nvidia/nemotron-3-nano-30b-a3b` | nemotron-3-nano-30b-a3b | 131072 | 131072 | minimal, low, medium, high | +| `moonshotai/kimi-k3` | Kimi K3 | 1048576 | 131072 | minimal, low, medium, high | +| `nvidia/cosmos-reason2-8b` | Cosmos Reason2 8B | 131072 | 16384 | minimal, low, medium, high | +| `nvidia/llama-3.1-nemotron-70b-instruct` | Llama 3.1 Nemotron 70B Instruct | 128k | 8192 | — | +| `nvidia/llama-3.1-nemotron-ultra-253b-v1` | Llama 3.1 Nemotron Ultra 253B | 128k | 16384 | minimal, low, medium, high | | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | Nemotron 3 Nano Omni | 256k | 65536 | minimal, low, medium, high | | `nvidia/nemotron-3-super-120b-a12b` | Nemotron 3 Super | 262144 | 262144 | minimal, low, medium, high | | `nvidia/nemotron-3-ultra-550b-a55b` | Nemotron 3 Ultra 550B A55B | 1000k | 65536 | minimal, low, medium, high | -| `nvidia/nvidia-nemotron-nano-9b-v2` | nvidia-nemotron-nano-9b-v2 | 131072 | 131072 | minimal, low, medium, high | -| `openai/gpt-oss-120b` | GPT-OSS-120B | 128k | 8192 | minimal, low, medium, high | +| `nvidia/nemotron-3.5-lightning-30b-a3b` | Nemotron 3.5 Lightning 30B A3B | 262144 | 262144 | minimal, low, medium, high | | `openai/gpt-oss-20b` | GPT OSS 20B | 131072 | 32768 | minimal, low, medium, high | -| `stepfun-ai/step-3.5-flash` | Step 3.5 Flash | 256k | 16384 | minimal, low, medium, high | -| `stepfun-ai/step-3.7-flash` | Step 3.7 Flash | 256k | 16384 | minimal, low, medium, high | -| `z-ai/glm-5.2` | GLM-5.2 | 1000k | 131072 | minimal, low, medium, high | +| `poolside/laguna-xs-2.1` | Laguna XS 2.1 | 262144 | 16384 | minimal, low, medium, high | +| `z-ai/glm-5.3` | GLM-5.3 | 1000k | 131072 | minimal, low, medium, high | +| `z-ai/glm-5.3-flash` | GLM-5.3-Flash | 1000k | 131072 | minimal, low, medium, high | ## openai-codex | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | | `gpt-5.3-codex-spark` | GPT-5.3 Codex Spark | 128k | 128k | minimal, low, medium, high, xhigh | -| `gpt-5.4` | GPT-5.4 | 272k | 128k | minimal, low, medium, high, xhigh | -| `gpt-5.4-mini` | GPT-5.4 mini | 272k | 128k | minimal, low, medium, high, xhigh | | `gpt-5.5` | GPT-5.5 | 272k | 128k | minimal, low, medium, high, xhigh | | `gpt-5.6-luna` | GPT-5.6 Luna | 272k | 128k | minimal, low, medium, high, xhigh, max | | `gpt-5.6-sol` | GPT-5.6 Sol | 272k | 128k | minimal, low, medium, high, xhigh, max | | `gpt-5.6-terra` | GPT-5.6 Terra | 272k | 128k | minimal, low, medium, high, xhigh, max | +| `gpt-6-astra` | GPT-6 Astra | 272k | 128k | minimal, low, medium, high, xhigh, max | +| `gpt-6-luna` | GPT-6 Luna | 272k | 128k | minimal, low, medium, high, xhigh, max | +| `gpt-6-sol` | GPT-6 Sol | 272k | 128k | minimal, low, medium, high, xhigh, max | ## openai @@ -612,6 +803,9 @@ connect a provider with `codegenie provider login ` (or env vars / the | `gpt-5.6-luna` | GPT-5.6 Luna | 272k | 128k | low, medium, high, xhigh, max | | `gpt-5.6-sol` | GPT-5.6 Sol | 272k | 128k | low, medium, high, xhigh, max | | `gpt-5.6-terra` | GPT-5.6 Terra | 272k | 128k | low, medium, high, xhigh, max | +| `gpt-6-astra` | GPT-6 Astra | 272k | 128k | low, medium, high, xhigh, max | +| `gpt-6-luna` | GPT-6 Luna | 272k | 128k | low, medium, high, xhigh, max | +| `gpt-6-sol` | GPT-6 Sol | 272k | 128k | low, medium, high, xhigh, max | | `gpt-realtime-2.1` | GPT-Realtime-2.1 | 128k | 32k | minimal, low, medium, high, xhigh | | `o1` | o1 | 200k | 100k | low, medium, high | | `o1-pro` | o1-pro | 200k | 100k | low, medium, high | @@ -624,22 +818,36 @@ connect a provider with `codegenie provider login ` (or env vars / the | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `deepseek-v4-flash` | DeepSeek V4 Flash | 1000k | 384k | high, max | -| `deepseek-v4-pro` | DeepSeek V4 Pro | 1000k | 384k | high, max | +| `deepseek-v4-flash` | DeepSeek V4 Flash | 1000k | 384k | low, high, max | +| `deepseek-v4-flash-vision-exp` | DeepSeek V4 Flash Vision Exp | 1000k | 384k | low, high, max | +| `deepseek-v4-pro` | DeepSeek V4 Pro (New) | 1000k | 384k | high, max | +| `deepseek-v4.1-flash` | DeepSeek V4.1 Flash | 1000k | 384k | low, high, max | | `glm-5.1` | GLM-5.1 | 202752 | 32768 | minimal, low, medium, high | | `glm-5.2` | GLM-5.2 | 1000k | 131072 | high, max | -| `grok-4.5` | Grok 4.5 | 500k | 500k | low, medium, high | -| `hy3` | Hy3 | 256k | 64k | low, high | +| `glm-5.3` | GLM-5.3 | 1000k | 131072 | low, high, max | +| `glm-5.3-flash` | GLM-5.3-Flash | 1000k | 131072 | low, high, max | +| `gpt-5.6-luna` | GPT-5.6 Luna | 1050k | 128k | low, medium, high, xhigh, max | +| `grok-4.6` | Grok 4.6 | 500k | 500k | low, medium, high, xhigh | +| `grok-4.7` | Grok 4.7 | 500k | 500k | low, medium, high, xhigh | +| `hy3` | Hy3 | 256k | 128k | low, high | +| `hy4-preview` | Hy4 preview | 1024k | 64k | high | | `kimi-k2.6` | Kimi K2.6 | 262144 | 65536 | high | | `kimi-k2.7-code` | Kimi K2.7 Code | 262144 | 262144 | minimal, low, medium, high | -| `kimi-k3` | Kimi K3 (2x usage) | 1048576 | 131072 | max | +| `kimi-k3` | Kimi K3 | 1048576 | 131072 | max | +| `longcat-2.0` | LongCat-2.0 | 1000k | 131072 | minimal, low, medium, high | | `mimo-v2.5` | MiMo V2.5 | 1000k | 128k | minimal, low, medium, high | | `mimo-v2.5-pro` | MiMo V2.5 Pro | 1048576 | 128k | minimal, low, medium, high | +| `mimo-v2.6-flash` | MiMo-V2.6-Flash | 1048576 | 131072 | minimal, low, medium, high | +| `mimo-v2.6-pro` | MiMo-V2.6-Pro | 1048576 | 131072 | minimal, low, medium, high | | `minimax-m2.7` | MiniMax-M2.7 | 204800 | 131072 | minimal, low, medium, high | | `minimax-m3` | MiniMax-M3 | 1000k | 131072 | minimal, low, medium, high | +| `muse-spark-1.2-contributor` | Muse Spark 1.2 Contributor | 1048576 | 131072 | minimal, low, medium, high, xhigh | +| `muse-spark-1.3-contributor` | Muse Spark 1.3 Contributor | 1048576 | 131072 | minimal, low, medium, high, xhigh | | `qwen3.6-plus` | Qwen3.6 Plus | 1000k | 65536 | minimal, low, medium, high | | `qwen3.7-max` | Qwen3.7 Max | 1000k | 65536 | minimal, low, medium, high | | `qwen3.7-plus` | Qwen3.7 Plus | 1000k | 65536 | minimal, low, medium, high | +| `qwen3.8-flash` | Qwen3.8 Flash | 1000k | 131072 | minimal, low, medium, high | +| `qwen3.8-max` | Qwen3.8 Max | 1000k | 131072 | low, medium, xhigh | ## opencode @@ -647,28 +855,34 @@ connect a provider with `codegenie provider login ` (or env vars / the | --- | --- | --- | --- | --- | | `big-pickle` | Big Pickle | 200k | 32k | minimal, low, medium, high | | `claude-fable-5` | Claude Fable 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-fable-5-1` | Claude Fable 5.1 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `claude-haiku-4-5` | Claude Haiku 4.5 | 200k | 64k | minimal, low, medium, high | -| `claude-opus-4-1` | Claude Opus 4.1 | 200k | 32k | minimal, low, medium, high | | `claude-opus-4-5` | Claude Opus 4.5 | 200k | 64k | minimal, low, medium, high | | `claude-opus-4-6` | Claude Opus 4.6 | 1000k | 128k | minimal, low, medium, high, max | | `claude-opus-4-7` | Claude Opus 4.7 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `claude-opus-4-8` | Claude Opus 4.8 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `claude-opus-5` | Claude Opus 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-opus-5-5` | Claude Opus 5.5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `claude-sonnet-4` | Claude Sonnet 4 | 200k | 64k | minimal, low, medium, high | | `claude-sonnet-4-5` | Claude Sonnet 4.5 | 200k | 64k | minimal, low, medium, high | | `claude-sonnet-4-6` | Claude Sonnet 4.6 | 1000k | 64k | minimal, low, medium, high, max | | `claude-sonnet-5` | Claude Sonnet 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `deepseek-v4-flash` | DeepSeek V4 Flash | 1000k | 384k | high, max | -| `deepseek-v4-flash-free` | DeepSeek V4 Flash Free | 200k | 128k | high, max | +| `deepseek-v4-flash` | DeepSeek V4 Flash | 1000k | 384k | low, high, max | +| `deepseek-v4-flash-vision-exp` | DeepSeek V4 Flash Vision Exp | 1000k | 384k | low, high, max | | `deepseek-v4-pro` | DeepSeek V4 Pro | 1000k | 384k | high, max | +| `deepseek-v4.1-flash` | DeepSeek V4.1 Flash | 1000k | 384k | low, high, max | | `gemini-3-flash` | Gemini 3 Flash | 1048576 | 65536 | minimal, low, medium, high | -| `gemini-3.1-pro` | Gemini 3.1 Pro Preview | 1048576 | 65536 | low, high | +| `gemini-3.1-pro` | Gemini 3.1 Pro Preview | 1048576 | 65536 | low, medium, high | | `gemini-3.5-flash` | Gemini 3.5 Flash | 1048576 | 65536 | minimal, low, medium, high | | `gemini-3.5-flash-lite` | Gemini 3.5 Flash Lite | 1048576 | 65536 | minimal, low, medium, high | | `gemini-3.6-flash` | Gemini 3.6 Flash | 1048576 | 65536 | minimal, low, medium, high | +| `gemini-3.7-flash` | Gemini 3.7 Flash | 1048576 | 65536 | low, medium, high | +| `gemini-3.8-flash` | Gemini 3.8 Flash | 1048576 | 65536 | low, medium, high | | `glm-5` | GLM-5 | 204800 | 131072 | minimal, low, medium, high | | `glm-5.1` | GLM-5.1 | 204800 | 131072 | minimal, low, medium, high | | `glm-5.2` | GLM-5.2 | 1000k | 131072 | high, max | +| `glm-5.3` | GLM-5.3 | 1000k | 131072 | low, high, max | +| `glm-5.3-flash` | GLM-5.3-Flash | 1000k | 131072 | low, high, max | | `gpt-5` | GPT-5 | 400k | 128k | minimal, low, medium, high | | `gpt-5-codex` | GPT-5 Codex | 400k | 128k | low, medium, high | | `gpt-5-nano` | GPT-5 Nano | 400k | 128k | minimal, low, medium, high | @@ -688,298 +902,418 @@ connect a provider with `codegenie provider login ` (or env vars / the | `gpt-5.6-luna` | GPT-5.6 Luna | 1050k | 128k | low, medium, high, xhigh, max | | `gpt-5.6-sol` | GPT-5.6 Sol | 1050k | 128k | low, medium, high, xhigh, max | | `gpt-5.6-terra` | GPT-5.6 Terra | 1050k | 128k | low, medium, high, xhigh, max | +| `gpt-6-astra` | GPT-6 Astra | 1050k | 128k | low, medium, high, xhigh, max | +| `gpt-6-luna` | GPT-6 Luna | 1050k | 128k | low, medium, high, xhigh, max | +| `gpt-6-sol` | GPT-6 Sol | 1050k | 128k | low, medium, high, xhigh, max | | `grok-4.5` | Grok 4.5 | 500k | 500k | low, medium, high | +| `grok-4.6` | Grok 4.6 | 500k | 500k | low, medium, high, xhigh | | `grok-build-0.1` | Grok Build 0.1 | 256k | 256k | high | | `kimi-k2.5` | Kimi K2.5 | 262144 | 65536 | minimal, low, medium, high | | `kimi-k2.6` | Kimi K2.6 | 262144 | 65536 | minimal, low, medium, high | | `kimi-k2.7-code` | Kimi K2.7 Code | 262144 | 262144 | minimal, low, medium, high | -| `laguna-s-2.1-free` | Laguna S 2.1 Free | 256k | 32k | low, medium, high | -| `ling-3.0-flash-free` | Ling-3.0-flash Free | 262144 | 32768 | low, medium, high | -| `mimo-v2.5-free` | MiMo V2.5 Free | 200k | 32k | minimal, low, medium, high | +| `kimi-k3` | Kimi K3 | 1048576 | 131072 | max | +| `ling-3.0-flash-fin-free` | Ling 3.0 Flash Fin Free | 262144 | 32768 | minimal, low, medium, high | +| `mimo-v2.6-flash-free` | MiMo-V2.6-Flash Free | 200k | 32k | minimal, low, medium, high | | `minimax-m2.5` | MiniMax-M2.5 | 204800 | 131072 | minimal, low, medium, high | | `minimax-m2.7` | MiniMax-M2.7 | 204800 | 131072 | minimal, low, medium, high | | `minimax-m3` | MiniMax-M3 | 512k | 128k | minimal, low, medium, high | +| `muse-spark-1.2` | Muse Spark 1.2 | 1048576 | 131072 | minimal, low, medium, high, xhigh | +| `muse-spark-1.2-contributor-free` | Muse Spark 1.2 Free | 1048576 | 131072 | minimal, low, medium, high, xhigh | +| `muse-spark-1.3` | Muse Spark 1.3 | 1048576 | 131072 | minimal, low, medium, high, xhigh | +| `muse-spark-1.3-contributor-free` | Muse Spark 1.3 Free | 1048576 | 131072 | minimal, low, medium, high, xhigh | | `nemotron-3-ultra-free` | Nemotron 3 Ultra Free | 1000k | 128k | minimal, low, medium, high | -| `north-mini-code-free` | North Mini Code Free | 256k | 64k | high | +| `nemotron-3.5-lightning-free` | Nemotron 3.5 Lightning Free | 262144 | 262144 | minimal, low, medium, high | | `qwen3.5-plus` | Qwen3.5 Plus | 262144 | 65536 | minimal, low, medium, high | | `qwen3.6-plus` | Qwen3.6 Plus | 262144 | 65536 | minimal, low, medium, high | +| `qwen3.8-flash` | Qwen3.8 Flash | 1000k | 131072 | minimal, low, medium, high | ## openrouter | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `~anthropic/claude-fable-latest` | Anthropic: Claude Fable Latest | 1000k | 128k | minimal, low, medium, high | -| `~anthropic/claude-haiku-latest` | Anthropic Claude Haiku Latest | 200k | 64k | minimal, low, medium, high | -| `~anthropic/claude-opus-latest` | Anthropic: Claude Opus Latest | 1000k | 128k | minimal, low, medium, high | -| `~anthropic/claude-sonnet-latest` | Anthropic Claude Sonnet Latest | 1000k | 128k | minimal, low, medium, high | -| `~google/gemini-flash-latest` | Google Gemini Flash Latest | 1048576 | 65536 | minimal, low, medium, high | -| `~google/gemini-pro-latest` | Google Gemini Pro Latest | 1048576 | 65536 | minimal, low, medium, high | -| `~moonshotai/kimi-latest` | MoonshotAI Kimi Latest | 1048576 | 131072 | minimal, low, medium, high | -| `~openai/gpt-latest` | OpenAI GPT Latest | 1050k | 128k | minimal, low, medium, high | -| `~openai/gpt-mini-latest` | OpenAI GPT Mini Latest | 400k | 128k | minimal, low, medium, high | -| `~x-ai/grok-latest` | xAI: Grok Latest | 500k | 4096 | minimal, low, medium, high | -| `ai21/jamba-large-1.7` | AI21: Jamba Large 1.7 | 256k | 4096 | — | -| `aion-labs/aion-2.0` | AionLabs: Aion-2.0 | 131072 | 32768 | minimal, low, medium, high | -| `aion-labs/aion-3.0` | AionLabs: Aion-3.0 | 131072 | 32768 | minimal, low, medium, high | -| `aion-labs/aion-3.0-mini` | AionLabs: Aion-3.0-Mini | 131072 | 32768 | minimal, low, medium, high | +| `~anthropic/claude-fable-latest` | Anthropic: Claude Fable Latest | 1000k | 128k | low, medium, high, xhigh, max | +| `~anthropic/claude-haiku-latest` | Anthropic: Claude Haiku Latest | 200k | 64k | minimal, low, medium, high | +| `~anthropic/claude-opus-latest` | Anthropic: Claude Opus Latest | 1000k | 128k | low, medium, high, xhigh, max | +| `~anthropic/claude-sonnet-latest` | Anthropic: Claude Sonnet Latest | 1000k | 128k | low, medium, high, xhigh, max | +| `~deepseek/deepseek-flash-latest` | DeepSeek: DeepSeek Flash Latest | 1048576 | 943718 | low, high, max | +| `~deepseek/deepseek-pro-latest` | DeepSeek: DeepSeek Pro Latest | 1048576 | 393216 | low, high, max | +| `~deepseek/deepseek-v4-flash-latest` | DeepSeek: DeepSeek V4 Flash Latest | 1048576 | 943718 | low, high, max | +| `~google/gemini-flash-latest` | Google: Gemini Flash Latest | 1048576 | 65536 | low, medium, high | +| `~google/gemini-pro-latest` | Google: Gemini Pro Latest | 1048576 | 65536 | low, medium, high | +| `~moonshotai/kimi-latest` | MoonshotAI: Kimi Latest | 1048576 | 131072 | low, high, max | +| `~openai/gpt-astra-latest` | OpenAI: GPT Astra Latest | 1050k | 128k | low, medium, high, xhigh, max | +| `~openai/gpt-luna-latest` | OpenAI: GPT Luna Latest | 1050k | 128k | low, medium, high, xhigh, max | +| `~openai/gpt-mini-latest` | OpenAI: GPT Mini Latest | 400k | 128k | low, medium, high, xhigh | +| `~openai/gpt-sol-latest` | OpenAI: GPT Sol Latest | 1050k | 128k | low, medium, high, xhigh, max | +| `~openai/gpt-terra-latest` | OpenAI: GPT Terra Latest | 1050k | 128k | low, medium, high, xhigh, max | +| `~x-ai/grok-latest` | xAI: Grok Latest | 500k | 450k | low, medium, high, xhigh | +| `~z-ai/glm-flash-latest` | Z.ai: GLM Flash Latest | 1048576 | 943718 | low, high, max | +| `~z-ai/glm-latest` | Z.ai: GLM Latest | 1048576 | 131072 | low, high, max | +| `aion-labs/aion-2.0` | AionLabs: Aion-2.0 | 1048576 | 32768 | minimal, low, medium, high | +| `aion-labs/aion-3.0` | AionLabs: Aion-3.0 | 1048576 | 32768 | minimal, low, medium, high | +| `aion-labs/aion-3.0-mini` | AionLabs: Aion-3.0-Mini | 1048576 | 32768 | minimal, low, medium, high | | `amazon/nova-2-lite-v1` | Amazon: Nova 2 Lite | 1000k | 65535 | minimal, low, medium, high | | `amazon/nova-lite-v1` | Amazon: Nova Lite 1.0 | 300k | 5120 | — | | `amazon/nova-micro-v1` | Amazon: Nova Micro 1.0 | 128k | 5120 | — | | `amazon/nova-premier-v1` | Amazon: Nova Premier 1.0 | 1000k | 32k | — | | `amazon/nova-pro-v1` | Amazon: Nova Pro 1.0 | 300k | 5120 | — | -| `anthropic/claude-fable-5` | Anthropic: Claude Fable 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `anthropic/claude-fable-5` | Anthropic: Claude Fable 5 | 1000k | 128k | low, medium, high, xhigh, max | +| `anthropic/claude-fable-5:batch` | Anthropic: Claude Fable 5 (batch) | 1000k | 128k | low, medium, high, xhigh, max | +| `anthropic/claude-fable-5.1` | Anthropic: Claude Fable 5.1 | 1000k | 128k | low, medium, high, xhigh, max | +| `anthropic/claude-fable-5.1:batch` | Anthropic: Claude Fable 5.1 (batch) | 1000k | 128k | low, medium, high, xhigh, max | | `anthropic/claude-haiku-4.5` | Anthropic: Claude Haiku 4.5 | 200k | 64k | minimal, low, medium, high | +| `anthropic/claude-haiku-4.5:batch` | Anthropic: Claude Haiku 4.5 (batch) | 200k | 64k | minimal, low, medium, high | | `anthropic/claude-opus-4.1` | Anthropic: Claude Opus 4.1 | 200k | 32k | minimal, low, medium, high | +| `anthropic/claude-opus-4.1:batch` | Anthropic: Claude Opus 4.1 (batch) | 200k | 32k | minimal, low, medium, high | | `anthropic/claude-opus-4.5` | Anthropic: Claude Opus 4.5 | 200k | 64k | minimal, low, medium, high | -| `anthropic/claude-opus-4.6` | Anthropic: Claude Opus 4.6 | 1000k | 128k | minimal, low, medium, high, max | -| `anthropic/claude-opus-4.7` | Anthropic: Claude Opus 4.7 | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `anthropic/claude-opus-4.7-fast` | Anthropic: Claude Opus 4.7 (Fast) | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `anthropic/claude-opus-4.8` | Anthropic: Claude Opus 4.8 | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `anthropic/claude-opus-4.8-fast` | Anthropic: Claude Opus 4.8 (Fast) | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `anthropic/claude-opus-5` | Claude Opus 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `anthropic/claude-opus-5-fast` | Claude Opus 5 (Fast) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `anthropic/claude-opus-4.5:batch` | Anthropic: Claude Opus 4.5 (batch) | 200k | 64k | minimal, low, medium, high | +| `anthropic/claude-opus-4.6` | Anthropic: Claude Opus 4.6 | 1000k | 128k | low, medium, high, max | +| `anthropic/claude-opus-4.6:batch` | Anthropic: Claude Opus 4.6 (batch) | 1000k | 128k | low, medium, high, max | +| `anthropic/claude-opus-4.7` | Anthropic: Claude Opus 4.7 | 1000k | 128k | low, medium, high, xhigh, max | +| `anthropic/claude-opus-4.7:batch` | Anthropic: Claude Opus 4.7 (batch) | 1000k | 128k | low, medium, high, xhigh, max | +| `anthropic/claude-opus-4.8` | Anthropic: Claude Opus 4.8 | 1000k | 128k | low, medium, high, xhigh, max | +| `anthropic/claude-opus-4.8:batch` | Anthropic: Claude Opus 4.8 (batch) | 1000k | 128k | low, medium, high, xhigh, max | +| `anthropic/claude-opus-5` | Anthropic: Claude Opus 5 | 1000k | 128k | low, medium, high, xhigh, max | +| `anthropic/claude-opus-5:batch` | Anthropic: Claude Opus 5 (batch) | 1000k | 128k | low, medium, high, xhigh, max | +| `anthropic/claude-opus-5.5` | Anthropic: Claude Opus 5.5 | 1000k | 128k | low, medium, high, xhigh, max | +| `anthropic/claude-opus-5.5:batch` | Anthropic: Claude Opus 5.5 (batch) | 1000k | 128k | low, medium, high, xhigh, max | | `anthropic/claude-sonnet-4.5` | Anthropic: Claude Sonnet 4.5 | 1000k | 64k | minimal, low, medium, high | -| `anthropic/claude-sonnet-4.6` | Anthropic: Claude Sonnet 4.6 | 1000k | 128k | minimal, low, medium, high, max | -| `anthropic/claude-sonnet-5` | Anthropic: Claude Sonnet 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | -| `arcee-ai/trinity-large-thinking` | Arcee AI: Trinity Large Thinking | 262144 | 262144 | minimal, low, medium, high | -| `arcee-ai/virtuoso-large` | Arcee AI: Virtuoso Large | 131072 | 64k | — | +| `anthropic/claude-sonnet-4.5:batch` | Anthropic: Claude Sonnet 4.5 (batch) | 1000k | 64k | minimal, low, medium, high | +| `anthropic/claude-sonnet-4.6` | Anthropic: Claude Sonnet 4.6 | 1000k | 128k | low, medium, high, max | +| `anthropic/claude-sonnet-4.6:batch` | Anthropic: Claude Sonnet 4.6 (batch) | 1000k | 128k | low, medium, high, max | +| `anthropic/claude-sonnet-5` | Anthropic: Claude Sonnet 5 | 1000k | 128k | low, medium, high, xhigh, max | +| `anthropic/claude-sonnet-5:batch` | Anthropic: Claude Sonnet 5 (batch) | 1000k | 128k | low, medium, high, xhigh, max | +| `arcee-ai/trinity-large-thinking` | Arcee AI: Trinity Large Thinking | 262144 | 80k | minimal, low, medium, high | | `auto` | Auto | 2000k | 30k | minimal, low, medium, high | | `bytedance-seed/seed-1.6` | ByteDance Seed: Seed 1.6 | 262144 | 32768 | minimal, low, medium, high | | `bytedance-seed/seed-1.6-flash` | ByteDance Seed: Seed 1.6 Flash | 262144 | 32768 | minimal, low, medium, high | +| `bytedance-seed/seed-2-1-turbo` | ByteDance Seed: Seed 2.1 Turbo | 262144 | 235929 | minimal, low, medium, high | +| `bytedance-seed/seed-2.0-code` | ByteDance Seed: Seed-2.0-Code | 262144 | 131072 | low, medium, high | | `bytedance-seed/seed-2.0-lite` | ByteDance Seed: Seed-2.0-Lite | 262144 | 131072 | minimal, low, medium, high | | `bytedance-seed/seed-2.0-mini` | ByteDance Seed: Seed-2.0-Mini | 262144 | 131072 | minimal, low, medium, high | | `cohere/command-r-08-2024` | Cohere: Command R (08-2024) | 128k | 4k | — | | `cohere/command-r-plus-08-2024` | Cohere: Command R+ (08-2024) | 128k | 4k | — | | `cohere/north-mini-code:free` | Cohere: North Mini Code (free) | 256k | 64k | minimal, low, medium, high | -| `deepseek/deepseek-chat` | DeepSeek: DeepSeek V3 | 128k | 16k | — | -| `deepseek/deepseek-chat-v3-0324` | DeepSeek: DeepSeek V3 0324 | 163840 | 65536 | — | +| `deepseek/deepseek-chat` | DeepSeek: DeepSeek V3 | 163840 | 16384 | — | +| `deepseek/deepseek-chat-v3-0324` | DeepSeek: DeepSeek V3 0324 | 163840 | 147456 | — | | `deepseek/deepseek-chat-v3.1` | DeepSeek: DeepSeek V3.1 | 163840 | 32768 | minimal, low, medium, high | | `deepseek/deepseek-r1` | DeepSeek: R1 | 64k | 16k | minimal, low, medium, high | | `deepseek/deepseek-r1-0528` | DeepSeek: R1 0528 | 163840 | 32768 | minimal, low, medium, high | | `deepseek/deepseek-v3.1-terminus` | DeepSeek: DeepSeek V3.1 Terminus | 131072 | 32768 | minimal, low, medium, high | | `deepseek/deepseek-v3.2` | DeepSeek: DeepSeek V3.2 | 163840 | 65536 | minimal, low, medium, high | | `deepseek/deepseek-v3.2-exp` | DeepSeek: DeepSeek V3.2 Exp | 163840 | 65536 | minimal, low, medium, high | -| `deepseek/deepseek-v4-flash` | DeepSeek: DeepSeek V4 Flash | 1048575 | 4096 | high, xhigh | -| `deepseek/deepseek-v4-pro` | DeepSeek: DeepSeek V4 Pro | 1048576 | 384k | high, xhigh | +| `deepseek/deepseek-v4-flash` | DeepSeek: DeepSeek V4 Flash 0423 | 1024k | 384k | high, xhigh | +| `deepseek/deepseek-v4-flash-0731` | DeepSeek: DeepSeek V4 Flash 0731 | 1048576 | 943718 | low, high, max | +| `deepseek/deepseek-v4-flash-vision-exp` | DeepSeek: DeepSeek V4 Flash Vision Exp | 1048576 | 943718 | low, high, max | +| `deepseek/deepseek-v4-pro` | DeepSeek: DeepSeek V4 Pro 0423 | 1024k | 384k | high, xhigh | +| `deepseek/deepseek-v4-pro-0813` | DeepSeek: DeepSeek V4 Pro 0813 | 1048576 | 384k | low, high, max | +| `deepseek/deepseek-v4.1-flash` | DeepSeek: DeepSeek V4.1 Flash | 1048576 | 384k | low, high, max | +| `deepseek/deepseek-v4.1-flash:batch` | DeepSeek: DeepSeek V4.1 Flash (batch) | 1048576 | 131072 | low, high, max | +| `dots-studio/dots-3-note-preview:free` | Dots Studio: Dots3-Note Preview (free) | 512k | 460800 | minimal, low, medium, high | | `google/gemini-2.5-flash` | Google: Gemini 2.5 Flash | 1048576 | 65535 | minimal, low, medium, high | | `google/gemini-2.5-flash-lite` | Google: Gemini 2.5 Flash Lite | 1048576 | 65535 | minimal, low, medium, high | +| `google/gemini-2.5-flash-lite:batch` | Google: Gemini 2.5 Flash Lite (batch) | 1048576 | 65535 | minimal, low, medium, high | +| `google/gemini-2.5-flash:batch` | Google: Gemini 2.5 Flash (batch) | 1048576 | 65535 | minimal, low, medium, high | | `google/gemini-2.5-pro` | Google: Gemini 2.5 Pro | 1048576 | 65536 | minimal, low, medium, high | | `google/gemini-2.5-pro-preview` | Google: Gemini 2.5 Pro Preview 06-05 | 1048576 | 65536 | minimal, low, medium, high | -| `google/gemini-2.5-pro-preview-05-06` | Google: Gemini 2.5 Pro Preview 05-06 | 1048576 | 65535 | minimal, low, medium, high | -| `google/gemini-3-flash-preview` | Google: Gemini 3 Flash Preview | 1048576 | 65535 | minimal, low, medium, high | +| `google/gemini-2.5-pro:batch` | Google: Gemini 2.5 Pro (batch) | 1048576 | 65536 | minimal, low, medium, high | +| `google/gemini-3-flash-preview` | Google: Gemini 3 Flash Preview | 1048576 | 65536 | minimal, low, medium, high | +| `google/gemini-3-flash-preview:batch` | Google: Gemini 3 Flash Preview (batch) | 1048576 | 65536 | minimal, low, medium, high | | `google/gemini-3-pro-image` | Google: Nano Banana Pro (Gemini 3 Pro Image) | 65536 | 32768 | minimal, low, medium, high | | `google/gemini-3.1-flash-lite` | Google: Gemini 3.1 Flash Lite | 1048576 | 65536 | minimal, low, medium, high | | `google/gemini-3.1-flash-lite-preview` | Google: Gemini 3.1 Flash Lite Preview | 1048576 | 65536 | minimal, low, medium, high | -| `google/gemini-3.1-pro-preview` | Google: Gemini 3.1 Pro Preview | 1048576 | 65536 | minimal, low, medium, high | -| `google/gemini-3.1-pro-preview-customtools` | Google: Gemini 3.1 Pro Preview Custom Tools | 1048576 | 65536 | minimal, low, medium, high | +| `google/gemini-3.1-flash-lite:batch` | Google: Gemini 3.1 Flash Lite (batch) | 1048576 | 65536 | minimal, low, medium, high | +| `google/gemini-3.1-pro-preview` | Google: Gemini 3.1 Pro Preview | 1048576 | 65536 | low, medium, high | +| `google/gemini-3.1-pro-preview-customtools` | Google: Gemini 3.1 Pro Preview Custom Tools | 1048576 | 65536 | low, medium, high | +| `google/gemini-3.1-pro-preview:batch` | Google: Gemini 3.1 Pro Preview (batch) | 1048576 | 65536 | low, medium, high | | `google/gemini-3.5-flash` | Google: Gemini 3.5 Flash | 1048576 | 65536 | minimal, low, medium, high | | `google/gemini-3.5-flash-lite` | Google: Gemini 3.5 Flash Lite | 1048576 | 65536 | minimal, low, medium, high | +| `google/gemini-3.5-flash-lite:batch` | Google: Gemini 3.5 Flash Lite (batch) | 1048576 | 65536 | minimal, low, medium, high | +| `google/gemini-3.5-flash:batch` | Google: Gemini 3.5 Flash (batch) | 1048576 | 65536 | minimal, low, medium, high | | `google/gemini-3.6-flash` | Google: Gemini 3.6 Flash | 1048576 | 65536 | minimal, low, medium, high | +| `google/gemini-3.6-flash:batch` | Google: Gemini 3.6 Flash (batch) | 1048576 | 65536 | minimal, low, medium, high | +| `google/gemini-3.7-flash` | Google: Gemini 3.7 Flash | 1048576 | 65536 | low, medium, high | +| `google/gemini-3.7-flash:batch` | Google: Gemini 3.7 Flash (batch) | 1048576 | 65536 | low, medium, high | +| `google/gemini-3.8-flash` | Google: Gemini 3.8 Flash | 1048576 | 65536 | low, medium, high | +| `google/gemini-3.8-flash:batch` | Google: Gemini 3.8 Flash (batch) | 1048576 | 65536 | low, medium, high | | `google/gemma-3-12b-it` | Google: Gemma 3 12B | 131072 | 16384 | — | -| `google/gemma-3-27b-it` | Google: Gemma 3 27B | 131072 | 131072 | — | -| `google/gemma-4-26b-a4b-it` | Google: Gemma 4 26B A4B | 262144 | 262144 | minimal, low, medium, high | -| `google/gemma-4-26b-a4b-it:free` | Google: Gemma 4 26B A4B (free) | 131072 | 32768 | minimal, low, medium, high | -| `google/gemma-4-31b-it` | Google: Gemma 4 31B | 262144 | 262144 | minimal, low, medium, high | +| `google/gemma-3-27b-it` | Google: Gemma 3 27B | 131072 | 117964 | — | +| `google/gemma-4-26b-a4b-it` | Google: Gemma 4 26B A4B | 262144 | 235929 | minimal, low, medium, high | +| `google/gemma-4-26b-a4b-it:free` | Google: Gemma 4 26B A4B (free) | 262144 | 32768 | minimal, low, medium, high | +| `google/gemma-4-31b-it` | Google: Gemma 4 31B | 262144 | 16384 | minimal, low, medium, high | | `google/gemma-4-31b-it:free` | Google: Gemma 4 31B (free) | 262144 | 32768 | minimal, low, medium, high | -| `ibm-granite/granite-4.1-8b` | IBM: Granite 4.1 8B | 131072 | 131072 | — | -| `inception/mercury-2` | Inception: Mercury 2 | 128k | 50k | minimal, low, medium, high | -| `inclusionai/ling-2.6-1t` | inclusionAI: Ling-2.6-1T | 262144 | 32768 | — | -| `inclusionai/ling-2.6-flash` | inclusionAI: Ling-2.6-flash | 262144 | 32768 | — | -| `inclusionai/ling-3.0-flash:free` | Ling-3.0-flash (free) | 262144 | 32768 | minimal, low, medium, high | -| `inclusionai/ring-2.6-1t` | inclusionAI: Ring-2.6-1T | 262144 | 65536 | minimal, low, medium, high | -| `kwaipilot/kat-coder-air-v2.5` | Kwaipilot: KAT-Coder-Air V2.5 | 256k | 80k | — | -| `kwaipilot/kat-coder-pro-v2` | Kwaipilot: KAT-Coder-Pro V2 | 256k | 80k | — | -| `kwaipilot/kat-coder-pro-v2.5` | Kwaipilot: KAT-Coder-Pro V2.5 | 256k | 80k | — | +| `ibm-granite/granite-4.2-8b` | IBM: Granite 4.2 8B | 131072 | 117964 | low, high | +| `inception/mercury-2` | Inception: Mercury 2 | 128k | 50k | low, medium, high | +| `inception/mercury-2.5` | Inception: Mercury 2.5 | 260k | 65536 | low, medium, high | +| `inclusionai/ling-3.0-flash` | inclusionAI: Ling 3.0 Flash | 262144 | 32768 | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash-fin` | inclusionAI: Ling 3.0 Flash Fin | 262144 | 235929 | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash-fin:free` | inclusionAI: Ling 3.0 Flash Fin (free) | 262144 | 32768 | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash-sante:free` | inclusionAI: Ling 3.0 Flash Sante (free) | 262144 | 32768 | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash-vl` | inclusionAI: Ling 3.0 Flash VL | 131072 | 32768 | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash-vl:free` | inclusionAI: Ling 3.0 Flash VL (free) | 262144 | 32768 | minimal, low, medium, high | +| `kwaipilot/kat-coder-pro-v2.5` | Kwaipilot: KAT-Coder-Pro V2.5 | 262144 | 235929 | — | +| `liquid/lfm-2.5-2.6b:free` | LiquidAI: LFM2.5-2.6B (free) | 65536 | 8192 | minimal, low, medium, high | | `meituan/longcat-2.0` | Meituan: LongCat 2.0 | 1048756 | 262144 | minimal, low, medium, high | | `meta-llama/llama-3.1-70b-instruct` | Meta: Llama 3.1 70B Instruct | 131072 | 16384 | — | -| `meta-llama/llama-3.1-8b-instruct` | Meta: Llama 3.1 8B Instruct | 131072 | 131072 | — | -| `meta-llama/llama-3.3-70b-instruct` | Meta: Llama 3.3 70B Instruct | 131072 | 128k | — | -| `meta-llama/llama-4-maverick` | Meta: Llama 4 Maverick | 1048576 | 16384 | — | +| `meta-llama/llama-3.1-8b-instruct` | Meta: Llama 3.1 8B Instruct | 131072 | 117964 | — | +| `meta-llama/llama-3.3-70b-instruct` | Meta: Llama 3.3 70B Instruct | 131072 | 16384 | — | +| `meta-llama/llama-4-maverick` | Meta: Llama 4 Maverick | 128k | 16384 | — | | `meta-llama/llama-4-scout` | Meta: Llama 4 Scout | 327680 | 16384 | — | -| `meta/muse-spark-1.1` | Meta: Muse Spark 1.1 | 1048576 | 4096 | minimal, low, medium, high | +| `meta/muse-glimmer-30b` | Meta: Muse Glimmer 30B | 131072 | 16384 | low, medium, high, xhigh | +| `meta/muse-spark-1.1` | Meta: Muse Spark 1.1 | 1048576 | 943718 | minimal, low, medium, high, xhigh | +| `meta/muse-spark-1.2` | Meta: Muse Spark 1.2 | 1048576 | 943718 | minimal, low, medium, high, xhigh | +| `meta/muse-spark-1.2-contributor` | Meta: Muse Spark 1.2 Contributor | 1048576 | 943718 | minimal, low, medium, high, xhigh | +| `meta/muse-spark-1.3` | Meta: Muse Spark 1.3 | 1048576 | 943718 | minimal, low, medium, high, xhigh, max | +| `meta/muse-spark-1.3-contributor` | Meta: Muse Spark 1.3 Contributor | 1048576 | 943718 | minimal, low, medium, high, xhigh, max | | `minimax/minimax-m1` | MiniMax: MiniMax M1 | 1000k | 40k | minimal, low, medium, high | | `minimax/minimax-m2` | MiniMax: MiniMax M2 | 204800 | 131072 | minimal, low, medium, high | | `minimax/minimax-m2.1` | MiniMax: MiniMax M2.1 | 204800 | 131072 | minimal, low, medium, high | -| `minimax/minimax-m2.5` | MiniMax: MiniMax M2.5 | 196608 | 196608 | minimal, low, medium, high | -| `minimax/minimax-m2.7` | MiniMax: MiniMax M2.7 | 196608 | 131072 | minimal, low, medium, high | +| `minimax/minimax-m2.5` | MiniMax: MiniMax M2.5 | 200k | 128k | minimal, low, medium, high | +| `minimax/minimax-m2.7` | MiniMax: MiniMax M2.7 | 204800 | 131072 | minimal, low, medium, high | | `minimax/minimax-m3` | MiniMax: MiniMax M3 | 524288 | 512k | minimal, low, medium, high | -| `mistralai/codestral-2508` | Mistral: Codestral 2508 | 256k | 4096 | — | -| `mistralai/devstral-2512` | Mistral: Devstral 2 2512 | 262144 | 4096 | — | -| `mistralai/ministral-14b-2512` | Mistral: Ministral 3 14B 2512 | 262144 | 4096 | — | -| `mistralai/ministral-3b-2512` | Mistral: Ministral 3 3B 2512 | 131072 | 4096 | — | -| `mistralai/ministral-8b-2512` | Mistral: Ministral 3 8B 2512 | 262144 | 4096 | — | -| `mistralai/mistral-large` | Mistral Large | 128k | 4096 | — | -| `mistralai/mistral-large-2407` | Mistral Large 2407 | 131072 | 4096 | — | -| `mistralai/mistral-large-2512` | Mistral: Mistral Large 3 2512 | 262144 | 4096 | — | -| `mistralai/mistral-medium-3` | Mistral: Mistral Medium 3 | 131072 | 4096 | — | -| `mistralai/mistral-medium-3-5` | Mistral: Mistral Medium 3.5 | 262144 | 4096 | minimal, low, medium, high | -| `mistralai/mistral-medium-3.1` | Mistral: Mistral Medium 3.1 | 131072 | 4096 | — | +| `mistralai/codestral-2508` | Mistral: Codestral 2508 | 256k | 204800 | — | +| `mistralai/codestral-2508:batch` | Mistral: Codestral 2508 (batch) | 256k | 204800 | — | +| `mistralai/devstral-2512` | Mistral: Devstral 2 2512 | 262144 | 209715 | — | +| `mistralai/ministral-14b-2512` | Mistral: Ministral 3 14B 2512 | 262144 | 209715 | — | +| `mistralai/ministral-3b-2512` | Mistral: Ministral 3 3B 2512 | 131072 | 104857 | — | +| `mistralai/ministral-8b-2512` | Mistral: Ministral 3 8B 2512 | 262144 | 209715 | — | +| `mistralai/ministral-8b-2512:batch` | Mistral: Ministral 3 8B 2512 (batch) | 262144 | 209715 | — | +| `mistralai/mistral-large` | Mistral Large | 128k | 102400 | — | +| `mistralai/mistral-large-2407` | Mistral Large 2407 | 131072 | 104857 | — | +| `mistralai/mistral-large-2512:batch` | Mistral: Mistral Large 3 2512 (batch) | 262144 | 209715 | — | +| `mistralai/mistral-medium-3` | Mistral: Mistral Medium 3 | 131072 | 104857 | — | +| `mistralai/mistral-medium-3-5` | Mistral: Mistral Medium 3.5 | 262144 | 209715 | high | +| `mistralai/mistral-medium-3-5:batch` | Mistral: Mistral Medium 3.5 (batch) | 262144 | 209715 | high | +| `mistralai/mistral-medium-3.1` | Mistral: Mistral Medium 3.1 | 131072 | 104857 | — | +| `mistralai/mistral-medium-3.1:batch` | Mistral: Mistral Medium 3.1 (batch) | 131072 | 104857 | — | | `mistralai/mistral-nemo` | Mistral: Mistral Nemo | 131072 | 16384 | — | -| `mistralai/mistral-saba` | Mistral: Saba | 32768 | 4096 | — | -| `mistralai/mistral-small-2603` | Mistral: Mistral Small 4 | 262144 | 4096 | minimal, low, medium, high | -| `mistralai/mistral-small-3.2-24b-instruct` | Mistral: Mistral Small 3.2 24B | 131072 | 4096 | — | -| `mistralai/mixtral-8x22b-instruct` | Mistral: Mixtral 8x22B Instruct | 65536 | 4096 | — | -| `mistralai/voxtral-small-24b-2507` | Mistral: Voxtral Small 24B 2507 | 32k | 4096 | — | -| `moonshotai/kimi-k2` | MoonshotAI: Kimi K2 0711 | 131072 | 100352 | — | -| `moonshotai/kimi-k2-0905` | MoonshotAI: Kimi K2 0905 | 262144 | 100352 | — | -| `moonshotai/kimi-k2-thinking` | MoonshotAI: Kimi K2 Thinking | 262144 | 100352 | minimal, low, medium, high | +| `mistralai/mistral-saba` | Mistral: Saba | 32768 | 26214 | — | +| `mistralai/mistral-small-2603` | Mistral: Mistral Small 4 | 262144 | 209715 | high | +| `mistralai/mistral-small-2603:batch` | Mistral: Mistral Small 4 (batch) | 262144 | 209715 | high | +| `mistralai/mistral-small-3.1-24b-instruct` | Mistral: Mistral Small 3.1 24B | 128k | 102400 | — | +| `mistralai/mistral-small-3.2-24b-instruct` | Mistral: Mistral Small 3.2 24B | 256k | 16384 | — | +| `mistralai/mixtral-8x22b-instruct` | Mistral: Mixtral 8x22B Instruct | 65536 | 52428 | — | +| `mistralai/voxtral-small-24b-2507` | Mistral: Voxtral Small 24B 2507 | 32768 | 26214 | — | +| `moonshotai/kimi-k2` | MoonshotAI: Kimi K2 0711 | 131072 | 98304 | — | +| `moonshotai/kimi-k2-0905` | MoonshotAI: Kimi K2 0905 | 262144 | 98304 | — | +| `moonshotai/kimi-k2-thinking` | MoonshotAI: Kimi K2 Thinking | 262144 | 98304 | minimal, low, medium, high | | `moonshotai/kimi-k2.5` | MoonshotAI: Kimi K2.5 | 262144 | 4096 | minimal, low, medium, high | -| `moonshotai/kimi-k2.6` | MoonshotAI: Kimi K2.6 | 262144 | 262144 | minimal, low, medium, high | -| `moonshotai/kimi-k2.7-code` | MoonshotAI: Kimi K2.7 Code | 262144 | 262144 | minimal, low, medium, high | -| `moonshotai/kimi-k3` | MoonshotAI: Kimi K3 | 1048576 | 131072 | minimal, low, medium, high | -| `nex-agi/nex-n2-mini` | Nex AGI: Nex-N2-Mini | 262144 | 262144 | minimal, low, medium, high | -| `nex-agi/nex-n2-pro` | Nex AGI: Nex-N2-Pro | 262144 | 262144 | minimal, low, medium, high | -| `nvidia/nemotron-3-nano-30b-a3b` | NVIDIA: Nemotron 3 Nano 30B A3B | 262144 | 228k | minimal, low, medium, high | -| `nvidia/nemotron-3-nano-30b-a3b:free` | NVIDIA: Nemotron 3 Nano 30B A3B (free) | 256k | 4096 | minimal, low, medium, high | +| `moonshotai/kimi-k2.6` | MoonshotAI: Kimi K2.6 | 262144 | 235929 | minimal, low, medium, high | +| `moonshotai/kimi-k2.7-code` | MoonshotAI: Kimi K2.7 Code | 262144 | 235929 | minimal, low, medium, high | +| `moonshotai/kimi-k3` | MoonshotAI: Kimi K3 | 1048576 | 131072 | low, high, max | +| `moonshotai/kimi-k3:batch` | MoonshotAI: Kimi K3 (batch) | 1048576 | 16384 | low, high, max | +| `nex-agi/nex-n2.5-mini:free` | Nex AGI: Nex-N2.5-Mini (free) | 262144 | 235929 | medium, high | +| `nex-agi/nex-n2.5-pro` | Nex AGI: Nex-N2.5-Pro | 262144 | 235929 | medium, high | +| `nex-agi/nex-n2.5-pro:free` | Nex AGI: Nex-N2.5-Pro (free) | 262144 | 235929 | medium, high | +| `nvidia/nemotron-3-nano-30b-a3b` | NVIDIA: Nemotron 3 Nano 30B A3B | 262144 | 235929 | minimal, low, medium, high | | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free` | NVIDIA: Nemotron 3 Nano Omni (free) | 256k | 65536 | minimal, low, medium, high | -| `nvidia/nemotron-3-super-120b-a12b` | NVIDIA: Nemotron 3 Super | 262144 | 16384 | minimal, low, medium, high | -| `nvidia/nemotron-3-super-120b-a12b:free` | NVIDIA: Nemotron 3 Super (free) | 262144 | 262144 | minimal, low, medium, high | -| `nvidia/nemotron-3-ultra-550b-a55b` | NVIDIA: Nemotron 3 Ultra | 512288 | 4096 | minimal, low, medium, high | -| `nvidia/nemotron-3-ultra-550b-a55b:free` | NVIDIA: Nemotron 3 Ultra (free) | 1000k | 65536 | minimal, low, medium, high | -| `nvidia/nemotron-nano-12b-v2-vl:free` | NVIDIA: Nemotron Nano 12B 2 VL (free) | 128k | 128k | minimal, low, medium, high | -| `nvidia/nemotron-nano-9b-v2:free` | NVIDIA: Nemotron Nano 9B V2 (free) | 128k | 4096 | minimal, low, medium, high | +| `nvidia/nemotron-3-super-120b-a12b` | NVIDIA: Nemotron 3 Super | 262144 | 235929 | low, medium | +| `nvidia/nemotron-3-super-120b-a12b:free` | NVIDIA: Nemotron 3 Super (free) | 262144 | 235929 | low, medium | +| `nvidia/nemotron-3-ultra-550b-a55b` | NVIDIA: Nemotron 3 Ultra | 202800 | 182520 | medium, high | +| `nvidia/nemotron-3-ultra-550b-a55b:free` | NVIDIA: Nemotron 3 Ultra (free) | 1000k | 65536 | medium, high | +| `nvidia/nemotron-3.5-lightning` | NVIDIA: Nemotron 3.5 Lightning | 262144 | 235929 | minimal, low, medium, high | +| `nvidia/nemotron-3.5-lightning:free` | NVIDIA: Nemotron 3.5 Lightning (free) | 1000k | 65536 | minimal, low, medium, high | | `openai/gpt-3.5-turbo` | OpenAI: GPT-3.5 Turbo | 16385 | 4096 | — | -| `openai/gpt-3.5-turbo-0613` | OpenAI: GPT-3.5 Turbo (older v0613) | 4095 | 4096 | — | +| `openai/gpt-3.5-turbo-0613` | OpenAI: GPT-3.5 Turbo (older v0613) | 4095 | 3685 | — | | `openai/gpt-3.5-turbo-16k` | OpenAI: GPT-3.5 Turbo 16k | 16385 | 4096 | — | +| `openai/gpt-3.5-turbo:batch` | OpenAI: GPT-3.5 Turbo (batch) | 16385 | 4096 | — | | `openai/gpt-4` | OpenAI: GPT-4 | 8191 | 4096 | — | | `openai/gpt-4-turbo` | OpenAI: GPT-4 Turbo | 128k | 4096 | — | -| `openai/gpt-4-turbo-preview` | OpenAI: GPT-4 Turbo Preview | 128k | 4096 | — | +| `openai/gpt-4-turbo:batch` | OpenAI: GPT-4 Turbo (batch) | 128k | 4096 | — | | `openai/gpt-4.1` | OpenAI: GPT-4.1 | 1047576 | 32768 | — | | `openai/gpt-4.1-mini` | OpenAI: GPT-4.1 Mini | 1047576 | 32768 | — | +| `openai/gpt-4.1-mini:batch` | OpenAI: GPT-4.1 Mini (batch) | 1047576 | 32768 | — | | `openai/gpt-4.1-nano` | OpenAI: GPT-4.1 Nano | 1047576 | 32768 | — | +| `openai/gpt-4.1-nano:batch` | OpenAI: GPT-4.1 Nano (batch) | 1047576 | 32768 | — | +| `openai/gpt-4.1:batch` | OpenAI: GPT-4.1 (batch) | 1047576 | 32768 | — | | `openai/gpt-4o` | OpenAI: GPT-4o | 128k | 16384 | — | | `openai/gpt-4o-2024-05-13` | OpenAI: GPT-4o (2024-05-13) | 128k | 4096 | — | | `openai/gpt-4o-2024-08-06` | OpenAI: GPT-4o (2024-08-06) | 128k | 16384 | — | | `openai/gpt-4o-2024-11-20` | OpenAI: GPT-4o (2024-11-20) | 128k | 16384 | — | | `openai/gpt-4o-mini` | OpenAI: GPT-4o-mini | 128k | 16384 | — | | `openai/gpt-4o-mini-2024-07-18` | OpenAI: GPT-4o-mini (2024-07-18) | 128k | 16384 | — | +| `openai/gpt-4o-mini:batch` | OpenAI: GPT-4o-mini (batch) | 128k | 16384 | — | +| `openai/gpt-4o:batch` | OpenAI: GPT-4o (batch) | 128k | 16384 | — | | `openai/gpt-5` | OpenAI: GPT-5 | 400k | 128k | minimal, low, medium, high | -| `openai/gpt-5-codex` | OpenAI: GPT-5 Codex | 400k | 128k | minimal, low, medium, high | | `openai/gpt-5-mini` | OpenAI: GPT-5 Mini | 400k | 128k | minimal, low, medium, high | +| `openai/gpt-5-mini:batch` | OpenAI: GPT-5 Mini (batch) | 400k | 128k | minimal, low, medium, high | | `openai/gpt-5-nano` | OpenAI: GPT-5 Nano | 400k | 128k | minimal, low, medium, high | -| `openai/gpt-5-pro` | OpenAI: GPT-5 Pro | 400k | 128k | minimal, low, medium, high | -| `openai/gpt-5.1` | OpenAI: GPT-5.1 | 400k | 128k | minimal, low, medium, high | -| `openai/gpt-5.1-chat` | OpenAI: GPT-5.1 Chat | 128k | 16384 | — | -| `openai/gpt-5.1-codex` | OpenAI: GPT-5.1-Codex | 400k | 128k | minimal, low, medium, high | -| `openai/gpt-5.1-codex-max` | OpenAI: GPT-5.1-Codex-Max | 400k | 128k | minimal, low, medium, high | -| `openai/gpt-5.1-codex-mini` | OpenAI: GPT-5.1-Codex-Mini | 400k | 100k | minimal, low, medium, high | -| `openai/gpt-5.2` | OpenAI: GPT-5.2 | 400k | 128k | minimal, low, medium, high, xhigh | -| `openai/gpt-5.2-chat` | OpenAI: GPT-5.2 Chat | 128k | 16384 | — | -| `openai/gpt-5.2-codex` | OpenAI: GPT-5.2-Codex | 400k | 128k | minimal, low, medium, high, xhigh | -| `openai/gpt-5.2-pro` | OpenAI: GPT-5.2 Pro | 400k | 128k | minimal, low, medium, high, xhigh | -| `openai/gpt-5.3-chat` | OpenAI: GPT-5.3 Chat | 128k | 16384 | — | -| `openai/gpt-5.3-codex` | OpenAI: GPT-5.3-Codex | 400k | 128k | minimal, low, medium, high, xhigh | -| `openai/gpt-5.4` | OpenAI: GPT-5.4 | 1050k | 128k | minimal, low, medium, high, xhigh | -| `openai/gpt-5.4-mini` | OpenAI: GPT-5.4 Mini | 400k | 128k | minimal, low, medium, high, xhigh | -| `openai/gpt-5.4-nano` | OpenAI: GPT-5.4 Nano | 400k | 128k | minimal, low, medium, high, xhigh | -| `openai/gpt-5.4-pro` | OpenAI: GPT-5.4 Pro | 1050k | 128k | minimal, low, medium, high, xhigh | -| `openai/gpt-5.5` | OpenAI: GPT-5.5 | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-5-nano:batch` | OpenAI: GPT-5 Nano (batch) | 400k | 128k | minimal, low, medium, high | +| `openai/gpt-5-pro` | OpenAI: GPT-5 Pro | 400k | 128k | high | +| `openai/gpt-5-pro:batch` | OpenAI: GPT-5 Pro (batch) | 400k | 128k | high | +| `openai/gpt-5:batch` | OpenAI: GPT-5 (batch) | 400k | 128k | minimal, low, medium, high | +| `openai/gpt-5.1` | OpenAI: GPT-5.1 | 400k | 128k | low, medium, high | +| `openai/gpt-5.1-codex` | OpenAI: GPT-5.1-Codex | 400k | 128k | low, medium, high | +| `openai/gpt-5.1-codex-max` | OpenAI: GPT-5.1-Codex-Max | 400k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.1-codex-mini` | OpenAI: GPT-5.1-Codex-Mini | 400k | 128k | low, medium, high | +| `openai/gpt-5.1:batch` | OpenAI: GPT-5.1 (batch) | 400k | 128k | low, medium, high | +| `openai/gpt-5.2` | OpenAI: GPT-5.2 | 400k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.2-chat` | OpenAI: GPT-5.2 Chat | 128k | 32k | — | +| `openai/gpt-5.2-codex` | OpenAI: GPT-5.2-Codex | 400k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.2-pro` | OpenAI: GPT-5.2 Pro | 400k | 128k | medium, high, xhigh | +| `openai/gpt-5.2-pro:batch` | OpenAI: GPT-5.2 Pro (batch) | 400k | 128k | medium, high, xhigh | +| `openai/gpt-5.2:batch` | OpenAI: GPT-5.2 (batch) | 400k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.3-codex` | OpenAI: GPT-5.3-Codex | 400k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.4` | OpenAI: GPT-5.4 | 1050k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.4-mini` | OpenAI: GPT-5.4 Mini | 400k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.4-mini:batch` | OpenAI: GPT-5.4 Mini (batch) | 400k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.4-nano` | OpenAI: GPT-5.4 Nano | 400k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.4-nano:batch` | OpenAI: GPT-5.4 Nano (batch) | 400k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.4-pro` | OpenAI: GPT-5.4 Pro | 1050k | 128k | medium, high, xhigh | +| `openai/gpt-5.4-pro:batch` | OpenAI: GPT-5.4 Pro (batch) | 1050k | 128k | medium, high, xhigh | +| `openai/gpt-5.4:batch` | OpenAI: GPT-5.4 (batch) | 1050k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.5` | OpenAI: GPT-5.5 | 1050k | 128k | low, medium, high, xhigh | | `openai/gpt-5.5-pro` | OpenAI: GPT-5.5 Pro | 1050k | 128k | medium, high, xhigh | -| `openai/gpt-5.6-luna` | OpenAI: GPT-5.6 Luna | 1050k | 128k | minimal, low, medium, high, xhigh, max | -| `openai/gpt-5.6-luna-pro` | OpenAI: GPT-5.6 Luna Pro | 1050k | 128k | minimal, low, medium, high, xhigh, max | -| `openai/gpt-5.6-sol` | OpenAI: GPT-5.6 Sol | 1050k | 128k | minimal, low, medium, high, xhigh, max | -| `openai/gpt-5.6-sol-pro` | OpenAI: GPT-5.6 Sol Pro | 1050k | 128k | minimal, low, medium, high, xhigh, max | -| `openai/gpt-5.6-terra` | OpenAI: GPT-5.6 Terra | 1050k | 128k | minimal, low, medium, high, xhigh, max | -| `openai/gpt-5.6-terra-pro` | OpenAI: GPT-5.6 Terra Pro | 1050k | 128k | minimal, low, medium, high, xhigh, max | +| `openai/gpt-5.5-pro:batch` | OpenAI: GPT-5.5 Pro (batch) | 1050k | 128k | medium, high, xhigh | +| `openai/gpt-5.5:batch` | OpenAI: GPT-5.5 (batch) | 1050k | 128k | low, medium, high, xhigh | +| `openai/gpt-5.6-luna` | OpenAI: GPT-5.6 Luna | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-5.6-luna-pro` | OpenAI: GPT-5.6 Luna Pro | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-5.6-luna-pro:batch` | OpenAI: GPT-5.6 Luna Pro (batch) | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-5.6-luna:batch` | OpenAI: GPT-5.6 Luna (batch) | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-5.6-sol` | OpenAI: GPT-5.6 Sol | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-5.6-sol-pro` | OpenAI: GPT-5.6 Sol Pro | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-5.6-sol-pro:batch` | OpenAI: GPT-5.6 Sol Pro (batch) | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-5.6-sol:batch` | OpenAI: GPT-5.6 Sol (batch) | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-5.6-terra` | OpenAI: GPT-5.6 Terra | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-5.6-terra-pro` | OpenAI: GPT-5.6 Terra Pro | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-5.6-terra-pro:batch` | OpenAI: GPT-5.6 Terra Pro (batch) | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-5.6-terra:batch` | OpenAI: GPT-5.6 Terra (batch) | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-astra` | OpenAI: GPT-6 Astra | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-astra-pro` | OpenAI: GPT-6 Astra Pro | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-astra-pro:batch` | OpenAI: GPT-6 Astra Pro (batch) | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-astra:batch` | OpenAI: GPT-6 Astra (batch) | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-luna` | OpenAI: GPT-6 Luna | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-luna-pro` | OpenAI: GPT-6 Luna Pro | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-luna-pro:batch` | OpenAI: GPT-6 Luna Pro (batch) | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-luna:batch` | OpenAI: GPT-6 Luna (batch) | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-sol` | OpenAI: GPT-6 Sol | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-sol-pro` | OpenAI: GPT-6 Sol Pro | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-sol-pro:batch` | OpenAI: GPT-6 Sol Pro (batch) | 1050k | 128k | low, medium, high, xhigh, max | +| `openai/gpt-6-sol:batch` | OpenAI: GPT-6 Sol (batch) | 1050k | 128k | low, medium, high, xhigh, max | | `openai/gpt-audio` | OpenAI: GPT Audio | 128k | 16384 | — | | `openai/gpt-audio-mini` | OpenAI: GPT Audio Mini | 128k | 16384 | — | | `openai/gpt-chat-latest` | OpenAI: GPT Chat Latest | 400k | 128k | — | -| `openai/gpt-oss-120b` | OpenAI: gpt-oss-120b | 131072 | 131072 | minimal, low, medium, high | -| `openai/gpt-oss-20b` | OpenAI: gpt-oss-20b | 131072 | 131072 | minimal, low, medium, high | -| `openai/gpt-oss-20b:free` | OpenAI: gpt-oss-20b (free) | 131072 | 32768 | minimal, low, medium, high | +| `openai/gpt-oss-120b` | OpenAI: gpt-oss-120b | 131072 | 65536 | low, medium, high | +| `openai/gpt-oss-20b` | OpenAI: gpt-oss-20b | 131072 | 32768 | low, medium, high | +| `openai/gpt-oss-20b:batch` | OpenAI: gpt-oss-20b (batch) | 131072 | 117964 | low, medium, high | | `openai/gpt-oss-safeguard-20b` | OpenAI: gpt-oss-safeguard-20b | 131072 | 65536 | minimal, low, medium, high | | `openai/o1` | OpenAI: o1 | 200k | 100k | minimal, low, medium, high | | `openai/o3` | OpenAI: o3 | 200k | 100k | minimal, low, medium, high | -| `openai/o3-deep-research` | OpenAI: o3 Deep Research | 200k | 100k | minimal, low, medium, high | | `openai/o3-mini` | OpenAI: o3 Mini | 200k | 100k | minimal, low, medium, high | -| `openai/o3-mini-high` | OpenAI: o3 Mini High | 200k | 100k | minimal, low, medium, high | +| `openai/o3-mini-high` | OpenAI: o3 Mini High | 200k | 100k | high | +| `openai/o3-mini:batch` | OpenAI: o3 Mini (batch) | 200k | 100k | minimal, low, medium, high | | `openai/o3-pro` | OpenAI: o3 Pro | 200k | 100k | minimal, low, medium, high | +| `openai/o3:batch` | OpenAI: o3 (batch) | 200k | 100k | minimal, low, medium, high | | `openai/o4-mini` | OpenAI: o4 Mini | 200k | 100k | minimal, low, medium, high | -| `openai/o4-mini-deep-research` | OpenAI: o4 Mini Deep Research | 200k | 100k | minimal, low, medium, high | -| `openai/o4-mini-high` | OpenAI: o4 Mini High | 200k | 100k | minimal, low, medium, high | +| `openai/o4-mini-high` | OpenAI: o4 Mini High | 200k | 100k | high | +| `openai/o4-mini:batch` | OpenAI: o4 Mini (batch) | 200k | 100k | minimal, low, medium, high | | `openrouter/auto` | Auto Router | 2000k | 4096 | minimal, low, medium, high | | `openrouter/auto-beta` | Auto Router (Beta) | 2000k | 4096 | minimal, low, medium, high | | `openrouter/free` | Free Models Router | 200k | 4096 | minimal, low, medium, high | | `openrouter/fusion` | OpenRouter: Fusion | 1000k | 30k | minimal, low, medium, high | -| `poolside/laguna-m.1` | Poolside: Laguna M.1 | 262144 | 32768 | minimal, low, medium, high | -| `poolside/laguna-m.1:free` | Poolside: Laguna M.1 (free) | 262144 | 32768 | minimal, low, medium, high | | `poolside/laguna-s-2.1` | Poolside: Laguna S 2.1 | 1048576 | 131072 | minimal, low, medium, high | | `poolside/laguna-s-2.1:free` | Poolside: Laguna S 2.1 (free) | 262144 | 32768 | minimal, low, medium, high | | `poolside/laguna-xs-2.1` | Poolside: Laguna XS 2.1 | 262144 | 32768 | minimal, low, medium, high | | `poolside/laguna-xs-2.1:free` | Poolside: Laguna XS 2.1 (free) | 262144 | 32768 | minimal, low, medium, high | +| `prism-ml/ternary-bonsai-2-27b` | PrismML: Ternary Bonsai 2 27B | 262144 | 32768 | medium, xhigh | | `qwen/qwen-2.5-72b-instruct` | Qwen2.5 72B Instruct | 32768 | 16384 | — | -| `qwen/qwen-2.5-7b-instruct` | Qwen: Qwen2.5 7B Instruct | 32768 | 32768 | — | +| `qwen/qwen-2.5-7b-instruct` | Qwen: Qwen2.5 7B Instruct | 32768 | 29491 | — | | `qwen/qwen-plus` | Qwen: Qwen-Plus | 1000k | 32768 | — | | `qwen/qwen-plus-2025-07-28` | Qwen: Qwen Plus 0728 | 1000k | 32768 | — | -| `qwen/qwen-plus-2025-07-28:thinking` | Qwen: Qwen Plus 0728 (thinking) | 1000k | 32768 | minimal, low, medium, high | -| `qwen/qwen3-14b` | Qwen: Qwen3 14B | 131072 | 8192 | minimal, low, medium, high | +| `qwen/qwen3-14b` | Qwen: Qwen3 14B | 40960 | 16384 | minimal, low, medium, high | | `qwen/qwen3-235b-a22b` | Qwen: Qwen3 235B A22B | 131072 | 8192 | minimal, low, medium, high | -| `qwen/qwen3-235b-a22b-2507` | Qwen: Qwen3 235B A22B Instruct 2507 | 262144 | 16384 | — | -| `qwen/qwen3-235b-a22b-thinking-2507` | Qwen: Qwen3 235B A22B Thinking 2507 | 131072 | 32768 | minimal, low, medium, high | -| `qwen/qwen3-30b-a3b` | Qwen: Qwen3 30B A3B | 131072 | 8192 | minimal, low, medium, high | -| `qwen/qwen3-30b-a3b-instruct-2507` | Qwen: Qwen3 30B A3B Instruct 2507 | 262144 | 4096 | — | +| `qwen/qwen3-235b-a22b-2507` | Qwen: Qwen3 235B A22B Instruct 2507 | 262144 | 235929 | — | +| `qwen/qwen3-235b-a22b-thinking-2507` | Qwen: Qwen3 235B A22B Thinking 2507 | 131072 | 117964 | minimal, low, medium, high | +| `qwen/qwen3-30b-a3b` | Qwen: Qwen3 30B A3B | 40960 | 16384 | minimal, low, medium, high | +| `qwen/qwen3-30b-a3b-instruct-2507` | Qwen: Qwen3 30B A3B Instruct 2507 | 128k | 32k | — | | `qwen/qwen3-30b-a3b-thinking-2507` | Qwen: Qwen3 30B A3B Thinking 2507 | 81920 | 32768 | minimal, low, medium, high | | `qwen/qwen3-32b` | Qwen: Qwen3 32B | 40960 | 16384 | minimal, low, medium, high | | `qwen/qwen3-8b` | Qwen: Qwen3 8B | 131072 | 8192 | minimal, low, medium, high | | `qwen/qwen3-coder` | Qwen: Qwen3 Coder 480B A35B | 262144 | 65536 | — | -| `qwen/qwen3-coder-30b-a3b-instruct` | Qwen: Qwen3 Coder 30B A3B Instruct | 160k | 32768 | — | +| `qwen/qwen3-coder-30b-a3b-instruct` | Qwen: Qwen3 Coder 30B A3B Instruct | 262144 | 235929 | — | | `qwen/qwen3-coder-flash` | Qwen: Qwen3 Coder Flash | 1000k | 65536 | — | -| `qwen/qwen3-coder-next` | Qwen: Qwen3 Coder Next | 262144 | 262144 | — | +| `qwen/qwen3-coder-next` | Qwen: Qwen3 Coder Next | 262144 | 235929 | — | | `qwen/qwen3-coder-plus` | Qwen: Qwen3 Coder Plus | 1000k | 65536 | — | -| `qwen/qwen3-max` | Qwen: Qwen3 Max | 262144 | 32768 | — | -| `qwen/qwen3-max-thinking` | Qwen: Qwen3 Max Thinking | 262144 | 32768 | minimal, low, medium, high | -| `qwen/qwen3-next-80b-a3b-instruct` | Qwen: Qwen3 Next 80B A3B Instruct | 262144 | 262144 | — | -| `qwen/qwen3-next-80b-a3b-thinking` | Qwen: Qwen3 Next 80B A3B Thinking | 131072 | 32768 | minimal, low, medium, high | +| `qwen/qwen3-max` | Qwen: Qwen3 Max | 262144 | 65536 | — | +| `qwen/qwen3-max-thinking` | Qwen: Qwen3 Max Thinking | 262144 | 65536 | minimal, low, medium, high | +| `qwen/qwen3-next-80b-a3b-instruct` | Qwen: Qwen3 Next 80B A3B Instruct | 262144 | 16384 | — | +| `qwen/qwen3-next-80b-a3b-thinking` | Qwen: Qwen3 Next 80B A3B Thinking | 262144 | 235929 | minimal, low, medium, high | | `qwen/qwen3-vl-235b-a22b-instruct` | Qwen: Qwen3 VL 235B A22B Instruct | 131072 | 32768 | — | | `qwen/qwen3-vl-235b-a22b-thinking` | Qwen: Qwen3 VL 235B A22B Thinking | 131072 | 32768 | minimal, low, medium, high | -| `qwen/qwen3-vl-30b-a3b-instruct` | Qwen: Qwen3 VL 30B A3B Instruct | 262144 | 16384 | — | +| `qwen/qwen3-vl-30b-a3b-instruct` | Qwen: Qwen3 VL 30B A3B Instruct | 131072 | 32768 | — | | `qwen/qwen3-vl-30b-a3b-thinking` | Qwen: Qwen3 VL 30B A3B Thinking | 131072 | 32768 | minimal, low, medium, high | | `qwen/qwen3-vl-32b-instruct` | Qwen: Qwen3 VL 32B Instruct | 131072 | 32768 | — | | `qwen/qwen3-vl-8b-instruct` | Qwen: Qwen3 VL 8B Instruct | 131072 | 32768 | — | | `qwen/qwen3-vl-8b-thinking` | Qwen: Qwen3 VL 8B Thinking | 131072 | 32768 | minimal, low, medium, high | | `qwen/qwen3.5-122b-a10b` | Qwen: Qwen3.5-122B-A10B | 262144 | 65536 | minimal, low, medium, high | | `qwen/qwen3.5-27b` | Qwen: Qwen3.5-27B | 262144 | 65536 | minimal, low, medium, high | -| `qwen/qwen3.5-35b-a3b` | Qwen: Qwen3.5-35B-A3B | 262144 | 262144 | minimal, low, medium, high | -| `qwen/qwen3.5-397b-a17b` | Qwen: Qwen3.5 397B A17B | 262144 | 65536 | minimal, low, medium, high | -| `qwen/qwen3.5-9b` | Qwen: Qwen3.5-9B | 262144 | 262144 | minimal, low, medium, high | +| `qwen/qwen3.5-35b-a3b` | Qwen: Qwen3.5-35B-A3B | 256k | 16384 | minimal, low, medium, high | +| `qwen/qwen3.5-397b-a17b` | Qwen: Qwen3.5 397B A17B | 262144 | 235929 | minimal, low, medium, high | +| `qwen/qwen3.5-9b` | Qwen: Qwen3.5-9B | 256k | 32768 | minimal, low, medium, high | | `qwen/qwen3.5-flash-02-23` | Qwen: Qwen3.5-Flash | 1000k | 65536 | minimal, low, medium, high | | `qwen/qwen3.5-plus-02-15` | Qwen: Qwen3.5 Plus 2026-02-15 | 1000k | 65536 | minimal, low, medium, high | | `qwen/qwen3.5-plus-20260420` | Qwen: Qwen3.5 Plus 2026-04-20 | 1000k | 65536 | minimal, low, medium, high | -| `qwen/qwen3.6-27b` | Qwen: Qwen3.6 27B | 131072 | 131072 | minimal, low, medium, high | -| `qwen/qwen3.6-35b-a3b` | Qwen: Qwen3.6 35B A3B | 262144 | 262144 | minimal, low, medium, high | +| `qwen/qwen3.6-27b` | Qwen: Qwen3.6 27B | 262144 | 262140 | minimal, low, medium, high | +| `qwen/qwen3.6-35b-a3b` | Qwen: Qwen3.6 35B A3B | 262144 | 235929 | minimal, low, medium, high | | `qwen/qwen3.6-flash` | Qwen: Qwen3.6 Flash | 1000k | 65536 | minimal, low, medium, high | | `qwen/qwen3.6-max-preview` | Qwen: Qwen3.6 Max Preview | 262144 | 65536 | minimal, low, medium, high | | `qwen/qwen3.6-plus` | Qwen: Qwen3.6 Plus | 1000k | 65536 | minimal, low, medium, high | -| `qwen/qwen3.7-max` | Qwen: Qwen3.7 Max | 1000k | 65536 | minimal, low, medium, high | -| `qwen/qwen3.7-plus` | Qwen: Qwen3.7 Plus | 1000k | 65536 | minimal, low, medium, high | -| `rekaai/reka-edge` | Reka Edge | 16384 | 16384 | — | +| `qwen/qwen3.7-flash` | Qwen: Qwen3.7 Flash | 1000k | 65536 | minimal, low, medium, high | +| `qwen/qwen3.7-max` | Qwen: Qwen3.7 Max | 1000k | 131072 | minimal, low, medium, high | +| `qwen/qwen3.7-plus` | Qwen: Qwen3.7 Plus | 1000k | 131072 | minimal, low, medium, high | +| `qwen/qwen3.8-2.4t-a95b` | Qwen: Qwen3.8 2.4T A95B | 1000k | 131072 | low, medium, xhigh | +| `qwen/qwen3.8-27b` | Qwen: Qwen3.8 27B | 1000k | 131072 | low, medium, xhigh | +| `qwen/qwen3.8-27b:free` | Qwen: Qwen3.8 27B (free) | 262144 | 235929 | low, medium, xhigh | +| `qwen/qwen3.8-flash` | Qwen: Qwen3.8 Flash | 1000k | 131072 | minimal, low, medium, high | +| `qwen/qwen3.8-max-0902` | Qwen: Qwen3.8 Max (0902) | 1000k | 131072 | minimal, low, medium, high, xhigh | +| `qwen/qwen3.8-omni-flash` | Qwen: Qwen3.8 Omni Flash | 1000k | 131072 | minimal, low, medium, high | +| `rekaai/reka-edge` | Reka Edge | 16384 | 14745 | — | | `relace/relace-search` | Relace: Relace Search | 256k | 128k | — | -| `sakana/fugu-ultra` | Sakana: Fugu Ultra | 1000k | 128k | minimal, low, medium, high | +| `sakana/fugu-max` | Sakana: Fugu Max | 1000k | 128k | high, xhigh, max | +| `sakana/fugu-ultra` | Sakana: Fugu Ultra | 1000k | 128k | high, xhigh, max | +| `sakana/fugu-ultra-v2` | Sakana: Fugu Ultra v2 | 1000k | 128k | high, xhigh, max | +| `sakana/sakana-namazu` | Sakana: Sakana Namazu | 262144 | 65536 | high | | `sao10k/l3.1-euryale-70b` | Sao10K: Llama 3.1 Euryale 70B v2.2 | 131072 | 16384 | — | | `stepfun/step-3.5-flash` | StepFun: Step 3.5 Flash | 262144 | 65536 | minimal, low, medium, high | -| `stepfun/step-3.7-flash` | StepFun: Step 3.7 Flash | 256k | 256k | minimal, low, medium, high | -| `tencent/hy3` | Tencent: Hy3 | 262144 | 128k | minimal, low, medium, high | -| `tencent/hy3-preview` | Tencent: Hy3 preview | 262144 | 4096 | minimal, low, medium, high | -| `thedrummer/unslopnemo-12b` | TheDrummer: UnslopNemo 12B | 32768 | 32768 | — | -| `thinkingmachines/inkling` | Thinking Machines: Inkling | 524288 | 4096 | minimal, low, medium, high | -| `upstage/solar-pro-3` | Upstage: Solar Pro 3 | 128k | 4096 | minimal, low, medium, high | -| `x-ai/grok-4.20` | xAI: Grok 4.20 | 2000k | 4096 | minimal, low, medium, high | -| `x-ai/grok-4.3` | xAI: Grok 4.3 | 1000k | 4096 | minimal, low, medium, high | -| `x-ai/grok-4.5` | xAI: Grok 4.5 | 500k | 4096 | minimal, low, medium, high | -| `x-ai/grok-build-0.1` | xAI: Grok Build 0.1 | 256k | 4096 | minimal, low, medium, high | +| `stepfun/step-3.7-flash` | StepFun: Step 3.7 Flash | 256k | 230400 | low, medium, high | +| `tencent/hy3` | Tencent: Hy3 | 262144 | 128k | low, high | +| `tencent/hy3-preview` | Tencent: Hy3 preview | 262144 | 235929 | low, high | +| `tencent/hy4-preview` | Tencent: Hy4 preview | 1048576 | 64k | low, high | +| `thinkingmachines/inkling` | Thinking Machines: Inkling | 524288 | 471859 | minimal, low, medium, high, max | +| `thinkingmachines/inkling-small` | Thinking Machines: Inkling Small | 524288 | 262144 | minimal, low, medium, high, max | +| `thinkingmachines/inkling-small:free` | Thinking Machines: Inkling Small (free) | 1048576 | 262144 | minimal, low, medium, high, max | +| `thinkingmachines/inkling:free` | Thinking Machines: Inkling (free) | 1048576 | 262144 | minimal, low, medium, high, max | +| `unbiased/pareto` | Pareto | 262144 | 131072 | — | +| `upstage/solar-pro-3` | Upstage: Solar Pro 3 | 131072 | 117964 | minimal, low, medium, high | +| `upstage/solar-pro4` | Upstage: Solar Pro 4 | 524288 | 131072 | minimal, low, medium, high, xhigh, max | +| `x-ai/grok-4.20` | SpaceXAI: Grok 4.20 | 2000k | 1800k | minimal, low, medium, high | +| `x-ai/grok-4.3` | SpaceXAI: Grok 4.3 | 1000k | 900k | low, medium, high | +| `x-ai/grok-4.3:batch` | SpaceXAI: Grok 4.3 (batch) | 1000k | 900k | low, medium, high | +| `x-ai/grok-4.5` | SpaceXAI: Grok 4.5 | 500k | 450k | low, medium, high | +| `x-ai/grok-4.6` | SpaceXAI: Grok 4.6 | 500k | 450k | low, medium, high, xhigh | +| `x-ai/grok-4.7` | SpaceXAI: Grok 4.7 | 500k | 450k | low, medium, high, xhigh | +| `x-ai/grok-build-0.1` | SpaceXAI: Grok Build 0.1 | 256k | 230400 | minimal, low, medium, high | | `xiaomi/mimo-v2.5` | Xiaomi: MiMo-V2.5 | 1048576 | 131072 | minimal, low, medium, high | | `xiaomi/mimo-v2.5-pro` | Xiaomi: MiMo-V2.5-Pro | 1048576 | 131072 | minimal, low, medium, high | +| `xiaomi/mimo-v2.6-flash` | Xiaomi: MiMo-V2.6-Flash | 1048576 | 131072 | minimal, low, medium, high | +| `xiaomi/mimo-v2.6-pro` | Xiaomi: MiMo-V2.6-Pro | 1048576 | 131072 | minimal, low, medium, high | +| `xiaomi/mimo-v2.6-pro-ultraspeed` | Xiaomi: MiMo-V2.6-Pro-UltraSpeed | 1048576 | 131072 | minimal, low, medium, high | | `z-ai/glm-4.5` | Z.ai: GLM 4.5 | 131072 | 98304 | minimal, low, medium, high | | `z-ai/glm-4.5-air` | Z.ai: GLM 4.5 Air | 131072 | 98304 | minimal, low, medium, high | | `z-ai/glm-4.5v` | Z.ai: GLM 4.5V | 65536 | 16384 | minimal, low, medium, high | -| `z-ai/glm-4.6` | Z.ai: GLM 4.6 | 202752 | 131072 | minimal, low, medium, high | +| `z-ai/glm-4.6` | Z.ai: GLM 4.6 | 198k | 16384 | minimal, low, medium, high | | `z-ai/glm-4.6v` | Z.ai: GLM 4.6V | 131072 | 32768 | minimal, low, medium, high | | `z-ai/glm-4.7` | Z.ai: GLM 4.7 | 202752 | 131072 | minimal, low, medium, high | -| `z-ai/glm-4.7-flash` | Z.ai: GLM 4.7 Flash | 202752 | 16384 | minimal, low, medium, high | -| `z-ai/glm-5` | Z.ai: GLM 5 | 204800 | 131072 | minimal, low, medium, high | +| `z-ai/glm-4.7-flash` | Z.ai: GLM 4.7 Flash | 131072 | 117964 | minimal, low, medium, high | +| `z-ai/glm-5` | Z.ai: GLM 5 | 198k | 128k | minimal, low, medium, high | | `z-ai/glm-5-turbo` | Z.ai: GLM 5 Turbo | 202752 | 131072 | minimal, low, medium, high | | `z-ai/glm-5.1` | Z.ai: GLM 5.1 | 200k | 128k | minimal, low, medium, high | -| `z-ai/glm-5.2` | Z.ai: GLM 5.2 | 1048576 | 131072 | minimal, low, medium, high, xhigh | +| `z-ai/glm-5.2` | Z.ai: GLM 5.2 | 1048576 | 131072 | high, xhigh | +| `z-ai/glm-5.3` | Z.ai: GLM 5.3 | 1048576 | 131072 | low, high, max | +| `z-ai/glm-5.3-flash` | Z.ai: GLM 5.3 Flash | 1048576 | 943718 | low, high, max | +| `z-ai/glm-5.3-flash:batch` | Z.ai: GLM 5.3 Flash (batch) | 1048576 | 131072 | low, high, max | +| `z-ai/glm-5.3-flashx` | Z.ai: GLM 5.3 FlashX | 1048576 | 131072 | low, high, max | +| `z-ai/glm-5.3:batch` | Z.ai: GLM 5.3 (batch) | 1048576 | 131072 | low, high, max | | `z-ai/glm-5v-turbo` | Z.ai: GLM 5V Turbo | 202752 | 131072 | minimal, low, medium, high | ## qwen-token-plan-cn @@ -988,10 +1322,14 @@ connect a provider with `codegenie provider login ` (or env vars / the | --- | --- | --- | --- | --- | | `deepseek-v3.2` | DeepSeek V3.2 | 131072 | 65536 | minimal, low, medium, high | | `deepseek-v4-flash` | DeepSeek V4 Flash | 1000k | 384k | high, max | +| `deepseek-v4-flash-0731` | DeepSeek V4 Flash 0731 | 1000k | 384k | high, max | | `deepseek-v4-pro` | DeepSeek V4 Pro | 1000k | 384k | high, max | -| `glm-5` | GLM-5 | 202752 | 16384 | minimal, low, medium, high | -| `glm-5.1` | GLM-5.1 | 202752 | 128k | minimal, low, medium, high | -| `glm-5.2` | GLM-5.2 | 1000k | 131072 | minimal, low, medium, high | +| `deepseek-v4-pro-0813` | DeepSeek V4 Pro 0813 | 1000k | 384k | high, max | +| `deepseek-v4.1-flash` | DeepSeek V4.1 Flash | 1000k | 384k | low, high, max | +| `glm-5` | GLM-5 | 202752 | 16384 | high, max | +| `glm-5.1` | GLM-5.1 | 202752 | 128k | high, max | +| `glm-5.2` | GLM-5.2 | 1000k | 131072 | high, max | +| `glm-5.3` | GLM-5.3 | 1000k | 131072 | low, high, max | | `kimi-k2.5` | Kimi K2.5 | 262144 | 98304 | minimal, low, medium, high | | `kimi-k2.6` | Kimi K2.6 | 262144 | 262144 | minimal, low, medium, high | | `kimi-k2.7-code` | Kimi K2.7 Code | 262144 | 262144 | minimal, low, medium, high | @@ -1000,7 +1338,22 @@ connect a provider with `codegenie provider login ` (or env vars / the | `qwen3.6-plus` | Qwen3.6 Plus | 1000k | 65536 | minimal, low, medium, high | | `qwen3.7-max` | Qwen3.7 Max | 1000k | 131072 | minimal, low, medium, high | | `qwen3.7-plus` | Qwen3.7 Plus | 1000k | 65536 | minimal, low, medium, high | -| `qwen3.8-max-preview` | Qwen3.8 Max Preview | 1000k | 131072 | minimal, low, medium, high | +| `qwen3.8-flash` | Qwen3.8 Flash | 1000k | 131072 | low, medium, xhigh | +| `qwen3.8-max` | Qwen3.8 Max | 1000k | 131072 | low, medium, xhigh | + +## qwen-token-plan-individual + +| Model | Name | Context | Max output | Reasoning levels | +| --- | --- | --- | --- | --- | +| `deepseek-v4-flash-0731` | DeepSeek V4 Flash 0731 | 1000k | 384k | high, max | +| `deepseek-v4-pro` | DeepSeek V4 Pro | 1000k | 384k | high, max | +| `deepseek-v4-pro-0813` | DeepSeek V4 Pro 0813 | 1000k | 384k | high, max | +| `glm-5.2` | GLM-5.2 | 1000k | 131072 | high, max | +| `qwen3.6-flash` | Qwen3.6 Flash | 1000k | 65536 | minimal, low, medium, high | +| `qwen3.7-max` | Qwen3.7 Max | 1000k | 131072 | minimal, low, medium, high | +| `qwen3.7-plus` | Qwen3.7 Plus | 1000k | 65536 | minimal, low, medium, high | +| `qwen3.8-flash` | Qwen3.8 Flash | 1000k | 131072 | low, medium, xhigh | +| `qwen3.8-max` | Qwen3.8 Max | 1000k | 131072 | low, medium, xhigh | ## qwen-token-plan @@ -1008,10 +1361,14 @@ connect a provider with `codegenie provider login ` (or env vars / the | --- | --- | --- | --- | --- | | `deepseek-v3.2` | DeepSeek V3.2 | 131072 | 65536 | minimal, low, medium, high | | `deepseek-v4-flash` | DeepSeek V4 Flash | 1000k | 384k | high, max | +| `deepseek-v4-flash-0731` | DeepSeek V4 Flash 0731 | 1000k | 384k | high, max | | `deepseek-v4-pro` | DeepSeek V4 Pro | 1000k | 384k | high, max | -| `glm-5` | GLM-5 | 202752 | 16384 | minimal, low, medium, high | -| `glm-5.1` | GLM-5.1 | 202752 | 128k | minimal, low, medium, high | -| `glm-5.2` | GLM-5.2 | 1000k | 131072 | minimal, low, medium, high | +| `deepseek-v4-pro-0813` | DeepSeek V4 Pro 0813 | 1000k | 384k | high, max | +| `deepseek-v4.1-flash` | DeepSeek V4.1 Flash | 1000k | 384k | low, high, max | +| `glm-5` | GLM-5 | 202752 | 16384 | high, max | +| `glm-5.1` | GLM-5.1 | 202752 | 128k | high, max | +| `glm-5.2` | GLM-5.2 | 1000k | 131072 | high, max | +| `glm-5.3` | GLM-5.3 | 1000k | 131072 | low, high, max | | `kimi-k2.5` | Kimi K2.5 | 262144 | 98304 | minimal, low, medium, high | | `kimi-k2.6` | Kimi K2.6 | 262144 | 262144 | minimal, low, medium, high | | `kimi-k2.7-code` | Kimi K2.7 Code | 262144 | 262144 | minimal, low, medium, high | @@ -1020,19 +1377,59 @@ connect a provider with `codegenie provider login ` (or env vars / the | `qwen3.6-plus` | Qwen3.6 Plus | 1000k | 65536 | minimal, low, medium, high | | `qwen3.7-max` | Qwen3.7 Max | 1000k | 131072 | minimal, low, medium, high | | `qwen3.7-plus` | Qwen3.7 Plus | 1000k | 65536 | minimal, low, medium, high | -| `qwen3.8-max-preview` | Qwen3.8 Max Preview | 1000k | 131072 | minimal, low, medium, high | +| `qwen3.8-flash` | Qwen3.8 Flash | 1000k | 131072 | low, medium, xhigh | +| `qwen3.8-max` | Qwen3.8 Max | 1000k | 131072 | low, medium, xhigh | + +## radius + +| Model | Name | Context | Max output | Reasoning levels | +| --- | --- | --- | --- | --- | +| `balanced` | Balanced | 1048576 | 131072 | low, high, max | +| `cheap` | Cheap | 1000k | 384k | minimal, low, medium, high, xhigh, max | +| `claude-fable-5` | Claude Fable 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-fable-5-1` | Claude Fable 5.1 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-haiku-4-5` | Claude Haiku 4.5 | 200k | 64k | minimal, low, medium, high | +| `claude-opus-4-5` | Claude Opus 4.5 | 200k | 64k | minimal, low, medium, high | +| `claude-opus-4-8` | Claude Opus 4.8 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-opus-5` | Claude Opus 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `claude-opus-5-5` | Claude Opus 5.5 | 1000k | 128k | low, medium, high, xhigh, max | +| `claude-sonnet-4-5` | Claude Sonnet 4.5 | 1000k | 64k | minimal, low, medium, high | +| `claude-sonnet-5` | Claude Sonnet 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `deepseek-v4-flash` | DeepSeek V4.1 Flash | 1000k | 384k | minimal, low, medium, high, xhigh, max | +| `deepseek-v4-pro` | DeepSeek V4 Pro | 1000k | 262144 | high, max | +| `deepseek-v4.1-flash` | DeepSeek V4.1 Flash | 1000k | 384k | low, high, max | +| `glm-5.2` | GLM 5.2 | 432k | 131072 | high, max | +| `glm-5.3` | GLM-5.3 | 1048576 | 131072 | low, high, max | +| `glm-5.3-flash` | GLM-5.3 Flash | 1000k | 131072 | minimal, low, medium, high, xhigh, max | +| `gpt-5.3-codex` | GPT 5.3 Codex | 400k | 128k | low, medium, high, xhigh | +| `gpt-5.4` | GPT 5.4 | 272k | 128k | low, medium, high, xhigh | +| `gpt-5.4-mini` | GPT 5.4 Mini | 400k | 128k | low, medium, high, xhigh | +| `gpt-5.5` | GPT 5.5 | 272k | 128k | low, medium, high, xhigh | +| `gpt-5.6-luna` | GPT 5.6 Luna | 272k | 128k | low, medium, high, xhigh, max | +| `gpt-5.6-sol` | GPT 5.6 Sol | 272k | 128k | low, medium, high, xhigh, max | +| `gpt-5.6-terra` | GPT 5.6 Terra | 272k | 128k | low, medium, high, xhigh, max | +| `gpt-6-astra` | GPT 6 Astra | 272k | 128k | low, medium, high, xhigh, max | +| `gpt-6-luna` | GPT 6 Luna | 1050k | 128k | low, medium, high, xhigh, max | +| `gpt-6-sol` | GPT 6 Sol | 1050k | 128k | low, medium, high, xhigh, max | +| `kimi-k2.7-code` | Kimi K2.7 Code | 262k | 262k | minimal, low, medium, high | +| `kimi-k3` | Kimi K3 | 1048576 | 131072 | low, high, max | +| `precise` | Precise | 272k | 128k | low, medium, high, xhigh, max | ## together | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | +| `deepseek-ai/DeepSeek-V4-Flash-0731` | DeepSeek V4 Flash 0731 | 1000k | 384k | high | | `deepseek-ai/DeepSeek-V4-Pro` | DeepSeek V4 Pro | 512k | 384k | high | +| `deepseek-ai/DeepSeek-V4-Pro-0813` | DeepSeek V4 Pro 0813 | 1048576 | 384k | high | +| `deepseek-ai/DeepSeek-V4.1-Flash` | DeepSeek V4.1 Flash | 1048576 | 384k | high | | `google/gemma-4-31B-it` | Gemma 4 31B Instruct | 262144 | 131072 | high | | `meta-llama/Llama-3.3-70B-Instruct-Turbo` | Llama 3.3 70B | 131072 | 131072 | — | | `MiniMaxAI/MiniMax-M2.7` | MiniMax-M2.7 | 202752 | 131072 | high | | `MiniMaxAI/MiniMax-M3` | MiniMax-M3 | 524288 | 250k | high | | `moonshotai/Kimi-K2.6` | Kimi K2.6 | 262144 | 131k | high | | `moonshotai/Kimi-K2.7-Code` | Kimi K2.7 Code | 262144 | 131072 | high | +| `moonshotai/Kimi-K3` | Kimi K3 | 1048576 | 131072 | high | | `nvidia/nemotron-3-ultra-550b-a55b` | Nemotron 3 Ultra 550B A55B | 512300 | 512300 | high | | `openai/gpt-oss-120b` | GPT OSS 120B | 131072 | 131072 | low, medium, high | | `openai/gpt-oss-20b` | GPT OSS 20B | 131072 | 131072 | low, medium, high | @@ -1041,7 +1438,9 @@ connect a provider with `codegenie provider login ` (or env vars / the | `Qwen/Qwen3.6-Plus` | Qwen3.6 Plus | 1000k | 500k | high | | `Qwen/Qwen3.7-Max` | Qwen3.7 Max | 1000k | 500k | — | | `thinkingmachines/Inkling` | Inkling | 524288 | 131072 | high | -| `zai-org/GLM-5.2` | GLM-5.2 | 262144 | 164k | high | +| `zai-org/GLM-5.2` | GLM-5.2 | 512k | 164k | high | +| `zai-org/GLM-5.3` | GLM-5.3 | 1048576 | 262144 | high | +| `zai-org/GLM-5.3-Flash` | GLM-5.3-Flash | 1048575 | 400k | high | ## vercel-ai-gateway @@ -1060,8 +1459,8 @@ connect a provider with `codegenie provider login ` (or env vars / the | `alibaba/qwen3-max` | Qwen3 Max | 262144 | 32768 | — | | `alibaba/qwen3-max-preview` | Qwen3 Max Preview | 262144 | 32768 | — | | `alibaba/qwen3-max-thinking` | Qwen 3 Max Thinking | 256k | 65536 | minimal, low, medium, high | -| `alibaba/qwen3-next-80b-a3b-instruct` | Qwen3 Next 80B A3B Instruct | 131072 | 32768 | — | -| `alibaba/qwen3-next-80b-a3b-thinking` | Qwen3 Next 80B A3B Thinking | 131072 | 32768 | minimal, low, medium, high | +| `alibaba/qwen3-next-80b-a3b-instruct` | Qwen3 Next 80B A3B Instruct | 262114 | 262114 | — | +| `alibaba/qwen3-next-80b-a3b-thinking` | Qwen3 Next 80B A3B Thinking | 262144 | 262144 | minimal, low, medium, high | | `alibaba/qwen3-vl-235b-a22b-instruct` | Qwen3 VL 235B A22B Instruct | 131072 | 129024 | — | | `alibaba/qwen3-vl-instruct` | Qwen3 VL 235B A22B Instruct | 131072 | 129024 | — | | `alibaba/qwen3-vl-thinking` | Qwen3 VL 235B A22B Thinking | 131072 | 32768 | minimal, low, medium, high | @@ -1069,15 +1468,22 @@ connect a provider with `codegenie provider login ` (or env vars / the | `alibaba/qwen3.5-plus` | Qwen 3.5 Plus | 1000k | 64k | minimal, low, medium, high | | `alibaba/qwen3.6-27b` | Qwen 3.6 27B | 256k | 256k | minimal, low, medium, high | | `alibaba/qwen3.6-plus` | Qwen 3.6 Plus | 1000k | 64k | minimal, low, medium, high | +| `alibaba/qwen3.7-flash` | Qwen 3.7 Flash | 991k | 64k | minimal, low, medium, high | | `alibaba/qwen3.7-max` | Qwen 3.7 Max | 991k | 64k | minimal, low, medium, high | | `alibaba/qwen3.7-plus` | Qwen 3.7 Plus | 1000k | 64k | minimal, low, medium, high | +| `alibaba/qwen3.8-2.4t-a95b` | Qwen3.8 2.4T A95B | 262144 | 128k | minimal, low, medium, high | +| `alibaba/qwen3.8-27b` | Qwen3.8 27B | 1000k | 131072 | minimal, low, medium, high | +| `alibaba/qwen3.8-flash` | Qwen 3.8 Flash | 991k | 128k | minimal, low, medium, high | +| `alibaba/qwen3.8-max` | Qwen 3.8 Max | 262144 | 128k | minimal, low, medium, high | +| `alibaba/qwen3.8-max-0902` | Qwen3.8 Max 0902 | 991k | 128k | minimal, low, medium, high | +| `alibaba/qwen3.8-omni-flash` | Qwen 3.8 Omni Flash | 1000k | 131072 | minimal, low, medium, high | | `amazon/nova-2-lite` | Nova 2 Lite | 1000k | 1000k | minimal, low, medium, high | | `amazon/nova-lite` | Nova Lite | 300k | 8192 | — | | `amazon/nova-micro` | Nova Micro | 128k | 8192 | — | | `amazon/nova-pro` | Nova Pro | 300k | 8192 | — | | `anthropic/claude-fable-5` | Claude Fable 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `anthropic/claude-fable-5.1` | Claude Fable 5.1 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `anthropic/claude-haiku-4.5` | Claude Haiku 4.5 | 200k | 64k | minimal, low, medium, high | -| `anthropic/claude-opus-4.1` | Claude Opus 4.1 | 200k | 32k | minimal, low, medium, high | | `anthropic/claude-opus-4.5` | Claude Opus 4.5 | 200k | 64k | minimal, low, medium, high | | `anthropic/claude-opus-4.6` | Claude Opus 4.6 | 1000k | 128k | minimal, low, medium, high, max | | `anthropic/claude-opus-4.7` | Claude Opus 4.7 | 1000k | 128k | minimal, low, medium, high, xhigh, max | @@ -1085,48 +1491,62 @@ connect a provider with `codegenie provider login ` (or env vars / the | `anthropic/claude-opus-4.8-fast` | Claude Opus 4.8 (Fast) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `anthropic/claude-opus-5` | Claude Opus 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `anthropic/claude-opus-5-fast` | Claude Opus 5 (Fast) | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `anthropic/claude-opus-5.5` | Claude Opus 5.5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | +| `anthropic/claude-opus-5.5-fast` | Claude Opus 5.5 (Fast) | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `anthropic/claude-sonnet-4.5` | Claude Sonnet 4.5 | 1000k | 64k | minimal, low, medium, high | | `anthropic/claude-sonnet-4.6` | Claude Sonnet 4.6 | 1000k | 128k | minimal, low, medium, high, max | | `anthropic/claude-sonnet-5` | Claude Sonnet 5 | 1000k | 128k | minimal, low, medium, high, xhigh, max | | `arcee-ai/trinity-large-thinking` | Trinity Large Thinking | 262100 | 80k | minimal, low, medium, high | -| `arcee-ai/trinity-mini` | Trinity Mini | 131072 | 131072 | — | | `bytedance/seed-1.6` | Seed 1.6 | 256k | 32k | minimal, low, medium, high | | `bytedance/seed-1.8` | Bytedance Seed 1.8 | 256k | 64k | minimal, low, medium, high | +| `bytedance/seed-2.1-turbo` | Seed 2.1 Turbo | 262144 | 262144 | minimal, low, medium, high | | `cohere/command-a` | Command A | 256k | 8k | — | | `deepseek/deepseek-r1` | DeepSeek-R1 | 128k | 8192 | minimal, low, medium, high | -| `deepseek/deepseek-v3` | DeepSeek V3 0324 | 163840 | 163840 | — | | `deepseek/deepseek-v3.1` | DeepSeek V3.1 | 163840 | 128k | minimal, low, medium, high | | `deepseek/deepseek-v3.1-terminus` | DeepSeek V3.1 Terminus | 131072 | 65536 | minimal, low, medium, high | | `deepseek/deepseek-v3.2` | DeepSeek V3.2 | 128k | 8k | — | -| `deepseek/deepseek-v3.2-thinking` | DeepSeek V3.2 Thinking | 128k | 8k | minimal, low, medium, high | +| `deepseek/deepseek-v3.2-thinking` | DeepSeek V3.2 Thinking | 128k | 8k | — | | `deepseek/deepseek-v4-flash` | DeepSeek V4 Flash | 1000k | 384k | minimal, low, medium, high | +| `deepseek/deepseek-v4-flash-0731` | DeepSeek V4 Flash 0731 | 1000k | 384k | minimal, low, medium, high | +| `deepseek/deepseek-v4-flash-vision-exp` | DeepSeek V4 Flash Vision Exp | 1048576 | 1048576 | minimal, low, medium, high | | `deepseek/deepseek-v4-pro` | DeepSeek V4 Pro | 1000k | 384k | minimal, low, medium, high | +| `deepseek/deepseek-v4-pro-0813` | DeepSeek V4 Pro 0813 | 1000k | 384k | minimal, low, medium, high | +| `deepseek/deepseek-v4.1-flash` | DeepSeek V4.1 Flash | 1048576 | 32768 | minimal, low, medium, high | | `google/gemini-2.5-flash` | Gemini 2.5 Flash | 1000k | 65536 | minimal, low, medium, high | | `google/gemini-2.5-flash-lite` | Gemini 2.5 Flash Lite | 1048576 | 65536 | minimal, low, medium, high | | `google/gemini-2.5-pro` | Gemini 2.5 Pro | 1048576 | 65536 | minimal, low, medium, high | | `google/gemini-3-flash` | Gemini 3 Flash | 1000k | 65k | minimal, low, medium, high | -| `google/gemini-3-pro-preview` | Gemini 3 Pro Preview | 1000k | 64k | minimal, low, medium, high | | `google/gemini-3.1-flash-lite` | Gemini 3.1 Flash Lite | 1000k | 65k | minimal, low, medium, high | | `google/gemini-3.1-pro-preview` | Gemini 3.1 Pro Preview | 1000k | 64k | minimal, low, medium, high | | `google/gemini-3.5-flash` | Gemini 3.5 Flash | 1000k | 64k | minimal, low, medium, high | | `google/gemini-3.5-flash-lite` | Gemini 3.5 Flash Lite | 1000k | 65k | minimal, low, medium, high | | `google/gemini-3.6-flash` | Gemini 3.6 Flash | 1000k | 64k | minimal, low, medium, high | -| `google/gemma-4-26b-a4b-it` | Gemma 4 26B A4B IT | 262144 | 131072 | minimal, low, medium, high | +| `google/gemini-3.7-flash` | Gemini 3.7 Flash | 1000k | 65536 | minimal, low, medium, high | +| `google/gemini-3.8-flash` | Gemini 3.8 Flash | 1000k | 65536 | minimal, low, medium, high | +| `google/gemma-4-26b-a4b-it` | Google Gemma 4 26B A4B | 262144 | 131072 | minimal, low, medium, high | | `google/gemma-4-31b-it` | Gemma 4 31B IT | 262144 | 131072 | minimal, low, medium, high | | `inception/mercury-2` | Mercury 2 | 128k | 128k | minimal, low, medium, high | +| `inception/mercury-2.5` | Mercury 2.5 | 260k | 65536 | minimal, low, medium, high | | `inception/mercury-coder-small` | Mercury Coder Small Beta | 32k | 16384 | — | -| `inclusionai/ling-3.0-flash-free` | Ling 3.0 Flash | 256k | 256k | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash` | Ling 3.0 Flash | 256k | 32k | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash-fin` | Ling 3.0 Flash Fin | 256k | 32k | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash-fin-free` | Ling 3.0 Flash Fin (Free) | 256k | 32k | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash-sante` | Ling 3.0 Flash Sante | 256k | 32k | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash-sante-free` | Ling 3.0 Flash Sante (Free) | 256k | 32k | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash-vl` | Ling 3.0 Flash VL | 256k | 32k | minimal, low, medium, high | +| `inclusionai/ling-3.0-flash-vl-free` | Ling 3.0 Flash VL (Free) | 256k | 32k | minimal, low, medium, high | | `interfaze/interfaze-beta` | Interfaze Beta | 1000k | 32k | minimal, low, medium, high | -| `kwaipilot/kat-coder-air-v2.5` | Kat Coder Air V2.5 | 256k | 80k | minimal, low, medium, high | -| `kwaipilot/kat-coder-pro-v1` | KAT-Coder-Pro V1 | 256k | 32k | — | -| `kwaipilot/kat-coder-pro-v2` | Kat Coder Pro V2 | 256k | 256k | minimal, low, medium, high | -| `kwaipilot/kat-coder-pro-v2.5` | Kat Coder Pro V2.5 | 256k | 80k | minimal, low, medium, high | | `meta/llama-3.1-70b` | Llama 3.1 70B Instruct | 128k | 8192 | — | | `meta/llama-3.1-8b` | Llama 3.1 8B Instruct | 128k | 8192 | — | | `meta/llama-3.3-70b` | Llama 3.3 70B Instruct | 128k | 8192 | — | | `meta/llama-4-maverick` | Llama 4 Maverick 17B Instruct | 128k | 8192 | — | | `meta/llama-4-scout` | Llama 4 Scout 17B Instruct | 128k | 8192 | — | +| `meta/muse-glimmer-30b` | Muse Glimmer 30B | 131072 | 131072 | minimal, low, medium, high | | `meta/muse-spark-1.1` | Muse Spark 1.1 | 1048576 | 1048576 | minimal, low, medium, high | +| `meta/muse-spark-1.2` | Muse Spark 1.2 | 1048576 | 1048576 | minimal, low, medium, high | +| `meta/muse-spark-1.2-contributor` | Muse Spark 1.2 Contributor | 1048576 | 1048576 | minimal, low, medium, high | +| `meta/muse-spark-1.3` | Muse Spark 1.3 | 1048576 | 1048576 | minimal, low, medium, high | +| `meta/muse-spark-1.3-contributor` | Muse Spark 1.3 Contributor | 1048576 | 1048576 | minimal, low, medium, high | | `minimax/minimax-m2` | MiniMax M2 | 205k | 205k | minimal, low, medium, high | | `minimax/minimax-m2.1` | MiniMax M2.1 | 204800 | 131072 | minimal, low, medium, high | | `minimax/minimax-m2.1-lightning` | MiniMax M2.1 Lightning | 204800 | 131072 | minimal, low, medium, high | @@ -1134,171 +1554,207 @@ connect a provider with `codegenie provider login ` (or env vars / the | `minimax/minimax-m2.5-highspeed` | MiniMax M2.5 High Speed | 204800 | 131k | minimal, low, medium, high | | `minimax/minimax-m2.7` | MiniMax M2.7 | 204800 | 131k | minimal, low, medium, high | | `minimax/minimax-m2.7-highspeed` | MiniMax M2.7 High Speed | 204800 | 131100 | minimal, low, medium, high | -| `minimax/minimax-m3` | MiniMax M3 | 1000k | 1000k | minimal, low, medium, high | +| `minimax/minimax-m3` | MiniMax M3 | 512k | 512k | minimal, low, medium, high | | `mistral/codestral` | Mistral Codestral | 128k | 4k | — | -| `mistral/devstral-2` | Devstral 2 | 256k | 256k | — | -| `mistral/devstral-small-2` | Devstral Small 2 | 256k | 256k | — | -| `mistral/magistral-medium` | Magistral Medium 2509 | 128k | 64k | minimal, low, medium, high | -| `mistral/magistral-small` | Magistral Small 2509 | 128k | 64k | minimal, low, medium, high | -| `mistral/ministral-14b` | Ministral 14B | 256k | 256k | — | -| `mistral/ministral-3b` | Ministral 3B | 128k | 4k | — | -| `mistral/ministral-8b` | Ministral 8B | 128k | 4k | — | -| `mistral/mistral-large-3` | Mistral Large 3 | 256k | 256k | — | -| `mistral/mistral-medium` | Mistral Medium 3.1 | 128k | 64k | — | -| `mistral/mistral-medium-3.5` | Mistral Medium Latest | 256k | 256k | minimal, low, medium, high | -| `mistral/mistral-nemo` | Mistral Nemo 12B | 128k | 128k | — | -| `mistral/mistral-small` | Mistral Small | 32k | 4k | — | -| `mistral/pixtral-12b` | Pixtral 12B 2409 | 128k | 4k | — | +| `mistral/ministral-14b` | Ministral 14B | 262144 | 256k | — | +| `mistral/ministral-3b` | Ministral 3B | 131072 | 4k | — | +| `mistral/ministral-8b` | Ministral 8B | 262144 | 4k | — | +| `mistral/mistral-large-3` | Mistral Large 3 | 262144 | 256k | — | +| `mistral/mistral-medium-3.5` | Mistral Medium Latest | 262144 | 256k | minimal, low, medium, high | +| `mistral/mistral-nemo` | Mistral Nemo 12B | 60288 | 16k | — | +| `mistral/mistral-small` | Mistral Small | 262144 | 4k | — | +| `mixedbread/toast-1` | Toast 1 | 131k | 4k | — | | `moonshotai/kimi-k2` | Kimi K2 Instruct | 131072 | 131072 | — | | `moonshotai/kimi-k2-thinking` | Kimi K2 Thinking | 216144 | 216144 | minimal, low, medium, high | -| `moonshotai/kimi-k2.5` | Kimi K2.5 | 262114 | 262114 | minimal, low, medium, high | +| `moonshotai/kimi-k2.5` | Kimi K2.5 | 256k | 256k | minimal, low, medium, high | | `moonshotai/kimi-k2.6` | Kimi K2.6 | 262k | 262k | minimal, low, medium, high | | `moonshotai/kimi-k2.7-code` | Kimi K2.7 Code | 256k | 32768 | minimal, low, medium, high | | `moonshotai/kimi-k2.7-code-highspeed` | Kimi K2.7 Code High Speed | 262144 | 32768 | minimal, low, medium, high | | `moonshotai/kimi-k3` | Kimi K3 | 1000k | 131072 | minimal, low, medium, high | +| `moonshotai/kimi-k3-fast` | Kimi K3 Fast | 1000k | 131072 | minimal, low, medium, high | | `nvidia/nemotron-3-nano-30b-a3b` | Nemotron 3 Nano 30B A3B | 262144 | 262144 | minimal, low, medium, high | | `nvidia/nemotron-3-super-120b-a12b` | NVIDIA Nemotron 3 Super 120B A12B | 256k | 32k | minimal, low, medium, high | | `nvidia/nemotron-3-ultra-550b-a55b` | Nemotron 3 Ultra | 1000k | 65k | minimal, low, medium, high | +| `nvidia/nemotron-3.5-lightning` | Nemotron 3.5 Lightning 30B | 262144 | 131072 | minimal, low, medium, high | | `nvidia/nemotron-nano-12b-v2-vl` | Nvidia Nemotron Nano 12B V2 VL | 131072 | 131072 | minimal, low, medium, high | | `nvidia/nemotron-nano-9b-v2` | Nvidia Nemotron Nano 9B V2 | 131072 | 131072 | minimal, low, medium, high | | `openai/gpt-3.5-turbo` | GPT-3.5 Turbo | 16385 | 4096 | — | | `openai/gpt-4-turbo` | GPT-4 Turbo | 128k | 4096 | — | | `openai/gpt-4.1` | GPT-4.1 | 1047576 | 32768 | — | +| `openai/gpt-4.1-fast` | GPT-4.1 (Fast) | 1047576 | 32768 | — | | `openai/gpt-4.1-mini` | GPT-4.1 mini | 1047576 | 32768 | — | +| `openai/gpt-4.1-mini-fast` | GPT-4.1 mini (Fast) | 1047576 | 32768 | — | | `openai/gpt-4.1-nano` | GPT-4.1 nano | 1047576 | 32768 | — | +| `openai/gpt-4.1-nano-fast` | GPT-4.1 nano (Fast) | 1047576 | 32768 | — | | `openai/gpt-4o` | GPT-4o | 128k | 16384 | — | +| `openai/gpt-4o-fast` | GPT-4o (Fast) | 128k | 16384 | — | | `openai/gpt-4o-mini` | GPT-4o mini | 128k | 16384 | — | +| `openai/gpt-4o-mini-fast` | GPT-4o mini (Fast) | 128k | 16384 | — | | `openai/gpt-5` | GPT-5 | 400k | 128k | minimal, low, medium, high | | `openai/gpt-5-codex` | GPT-5-Codex | 400k | 128k | minimal, low, medium, high | +| `openai/gpt-5-fast` | GPT-5 (Fast) | 400k | 128k | minimal, low, medium, high | | `openai/gpt-5-mini` | GPT-5 mini | 400k | 128k | minimal, low, medium, high | +| `openai/gpt-5-mini-fast` | GPT-5 mini (Fast) | 400k | 128k | minimal, low, medium, high | | `openai/gpt-5-nano` | GPT-5 nano | 400k | 128k | minimal, low, medium, high | | `openai/gpt-5-pro` | GPT-5 pro | 400k | 272k | minimal, low, medium, high | | `openai/gpt-5.1-codex` | GPT-5.1-Codex | 400k | 128k | minimal, low, medium, high | | `openai/gpt-5.1-codex-max` | GPT 5.1 Codex Max | 400k | 128k | minimal, low, medium, high | | `openai/gpt-5.1-codex-mini` | GPT 5.1 Codex Mini | 400k | 128k | minimal, low, medium, high | -| `openai/gpt-5.1-instant` | GPT-5.1 Instant | 128k | 16384 | — | | `openai/gpt-5.1-thinking` | GPT 5.1 Thinking | 400k | 128k | minimal, low, medium, high | +| `openai/gpt-5.1-thinking-fast` | GPT 5.1 Thinking (Fast) | 400k | 128k | minimal, low, medium, high | | `openai/gpt-5.2` | GPT 5.2 | 400k | 128k | minimal, low, medium, high, xhigh | | `openai/gpt-5.2-codex` | GPT 5.2 Codex | 400k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-5.2-fast` | GPT 5.2 (Fast) | 400k | 128k | minimal, low, medium, high, xhigh | | `openai/gpt-5.2-pro` | GPT 5.2 | 400k | 128k | minimal, low, medium, high, xhigh | -| `openai/gpt-5.3-chat` | GPT-5.3 Chat | 128k | 16384 | — | | `openai/gpt-5.3-codex` | GPT 5.3 Codex | 400k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-5.3-codex-fast` | GPT 5.3 Codex (Fast) | 400k | 128k | minimal, low, medium, high, xhigh | | `openai/gpt-5.4` | GPT 5.4 | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-5.4-fast` | GPT 5.4 (Fast) | 1050k | 128k | minimal, low, medium, high, xhigh | | `openai/gpt-5.4-mini` | GPT 5.4 Mini | 400k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-5.4-mini-fast` | GPT 5.4 Mini (Fast) | 400k | 128k | minimal, low, medium, high, xhigh | | `openai/gpt-5.4-nano` | GPT 5.4 Nano | 400k | 128k | minimal, low, medium, high, xhigh | | `openai/gpt-5.4-pro` | GPT 5.4 Pro | 1050k | 128k | minimal, low, medium, high, xhigh | | `openai/gpt-5.5` | GPT 5.5 | 1000k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-5.5-fast` | GPT 5.5 (Fast) | 1000k | 128k | minimal, low, medium, high, xhigh | | `openai/gpt-5.5-pro` | GPT 5.5 Pro | 1000k | 128k | medium, high, xhigh | | `openai/gpt-5.6-luna` | GPT 5.6 Luna | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-5.6-luna-fast` | GPT 5.6 Luna (Fast) | 1050k | 128k | minimal, low, medium, high, xhigh | | `openai/gpt-5.6-sol` | GPT 5.6 Sol | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-5.6-sol-fast` | GPT 5.6 Sol (Fast) | 1050k | 128k | minimal, low, medium, high, xhigh | | `openai/gpt-5.6-terra` | GPT 5.6 Terra | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-5.6-terra-fast` | GPT 5.6 Terra (Fast) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-6-astra` | GPT-6 Astra | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-6-astra-fast` | GPT-6 Astra (Fast) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-6-luna` | GPT-6 Luna | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-6-luna-fast` | GPT-6 Luna (Fast) | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-6-sol` | GPT-6 Sol | 1050k | 128k | minimal, low, medium, high, xhigh | +| `openai/gpt-6-sol-fast` | GPT-6 Sol (Fast) | 1050k | 128k | minimal, low, medium, high, xhigh | | `openai/gpt-oss-120b` | GPT OSS 120B | 131072 | 131072 | minimal, low, medium, high | | `openai/gpt-oss-20b` | GPT OSS 20B | 131072 | 8192 | minimal, low, medium, high | -| `openai/gpt-oss-safeguard-20b` | GPT OSS Safeguard 20B | 131072 | 65536 | minimal, low, medium, high | +| `openai/gpt-oss-safeguard-120b` | GPT OSS Safeguard 120B | 128k | 16k | minimal, low, medium, high | +| `openai/gpt-oss-safeguard-20b` | GPT OSS Safeguard 20B | 128k | 16k | minimal, low, medium, high | | `openai/o1` | o1 | 200k | 100k | minimal, low, medium, high | | `openai/o3` | o3 | 200k | 100k | minimal, low, medium, high | -| `openai/o3-deep-research` | o3-deep-research | 200k | 100k | minimal, low, medium, high | +| `openai/o3-fast` | o3 (Fast) | 200k | 100k | minimal, low, medium, high | | `openai/o3-mini` | o3-mini | 200k | 100k | minimal, low, medium, high | | `openai/o3-pro` | o3 Pro | 200k | 100k | minimal, low, medium, high | | `openai/o4-mini` | o4-mini | 200k | 100k | minimal, low, medium, high | +| `openai/o4-mini-fast` | o4-mini (Fast) | 200k | 100k | minimal, low, medium, high | | `poolside/laguna-s-2.1` | Laguna S 2.1 | 1000k | 131072 | minimal, low, medium, high | | `poolside/laguna-s-2.1-free` | Laguna S 2.1 Free | 256k | 32768 | minimal, low, medium, high | +| `quiverai/arrow-2` | Arrow 2 | 131072 | 131072 | minimal, low, medium, high | +| `quiverai/arrow-2-telos` | Arrow 2 Telos | 131072 | 131072 | minimal, low, medium, high | +| `sakana/fugu-max` | Fugu Max | 1000k | 1000k | minimal, low, medium, high | | `sakana/fugu-ultra` | Fugu Ultra | 1000k | 1000k | minimal, low, medium, high | +| `sakana/fugu-ultra-v2` | Fugu Ultra v2 | 1000k | 1000k | minimal, low, medium, high | +| `sakana/namazu` | Sakana Namazu | 256k | 256k | minimal, low, medium, high | +| `spacexai/grok-4.1-fast-non-reasoning` | Grok 4.1 Fast Non-Reasoning | 1000k | 1000k | — | +| `spacexai/grok-4.1-fast-reasoning` | Grok 4.1 Fast Reasoning | 1000k | 1000k | minimal, low, medium, high | +| `spacexai/grok-4.20-multi-agent` | Grok 4.20 Multi-Agent | 2000k | 2000k | minimal, low, medium, high | +| `spacexai/grok-4.20-multi-agent-beta` | Grok 4.20 Multi Agent Beta | 2000k | 2000k | minimal, low, medium, high | +| `spacexai/grok-4.20-non-reasoning` | Grok 4.20 Non-Reasoning | 2000k | 2000k | — | +| `spacexai/grok-4.20-non-reasoning-beta` | Grok 4.20 Beta Non-Reasoning | 2000k | 2000k | — | +| `spacexai/grok-4.20-reasoning` | Grok 4.20 Reasoning | 2000k | 2000k | minimal, low, medium, high | +| `spacexai/grok-4.20-reasoning-beta` | Grok 4.20 Beta Reasoning | 2000k | 2000k | minimal, low, medium, high | +| `spacexai/grok-4.3` | Grok 4.3 | 1000k | 1000k | minimal, low, medium, high | +| `spacexai/grok-4.5` | Grok 4.5 | 500k | 500k | minimal, low, medium, high | +| `spacexai/grok-4.6` | Grok 4.6 | 500k | 500k | minimal, low, medium, high | +| `spacexai/grok-4.7` | Grok 4.7 | 500k | 500k | minimal, low, medium, high | +| `spacexai/grok-build-0.1` | Grok Build 0.1 | 256k | 256k | minimal, low, medium, high | | `stepfun/step-3.5-flash` | StepFun 3.5 Flash | 262114 | 262114 | minimal, low, medium, high | | `stepfun/step-3.7-flash` | Step 3.7 Flash | 256k | 256k | minimal, low, medium, high | | `tencent/hy3` | Hy3 | 262144 | 262144 | minimal, low, medium, high | +| `tencent/hy4-preview` | Tencent Hy4 Preview | 1024k | 64k | minimal, low, medium, high | | `thinkingmachines/inkling` | Inkling | 256k | 256k | minimal, low, medium, high | -| `xai/grok-4.1-fast-non-reasoning` | Grok 4.1 Fast Non-Reasoning | 1000k | 1000k | — | -| `xai/grok-4.1-fast-reasoning` | Grok 4.1 Fast Reasoning | 1000k | 1000k | minimal, low, medium, high | -| `xai/grok-4.20-multi-agent` | Grok 4.20 Multi-Agent | 2000k | 2000k | minimal, low, medium, high | -| `xai/grok-4.20-multi-agent-beta` | Grok 4.20 Multi Agent Beta | 2000k | 2000k | minimal, low, medium, high | -| `xai/grok-4.20-non-reasoning` | Grok 4.20 Non-Reasoning | 2000k | 2000k | — | -| `xai/grok-4.20-non-reasoning-beta` | Grok 4.20 Beta Non-Reasoning | 2000k | 2000k | — | -| `xai/grok-4.20-reasoning` | Grok 4.20 Reasoning | 2000k | 2000k | minimal, low, medium, high | -| `xai/grok-4.20-reasoning-beta` | Grok 4.20 Beta Reasoning | 2000k | 2000k | minimal, low, medium, high | -| `xai/grok-4.3` | Grok 4.3 | 1000k | 1000k | minimal, low, medium, high | -| `xai/grok-4.5` | Grok 4.5 | 500k | 500k | minimal, low, medium, high | -| `xai/grok-build-0.1` | Grok Build 0.1 | 256k | 256k | minimal, low, medium, high | +| `thinkingmachines/inkling-small` | Inkling Small | 1000k | 1000k | minimal, low, medium, high | | `xiaomi/mimo-v2.5` | MiMo M2.5 | 1050k | 131100 | minimal, low, medium, high | | `xiaomi/mimo-v2.5-pro` | MiMo V2.5 Pro | 1050k | 131k | minimal, low, medium, high | +| `xiaomi/mimo-v2.6-flash` | MiMo V2.6 Flash | 1048576 | 131072 | minimal, low, medium, high | +| `xiaomi/mimo-v2.6-pro` | MiMo V2.6 Pro | 1048576 | 131072 | minimal, low, medium, high | +| `xiaomi/mimo-v2.6-pro-ultraspeed` | MiMo V2.6 Pro UltraSpeed | 1048576 | 131072 | minimal, low, medium, high | | `zai/glm-4.5` | GLM 4.5 | 128k | 96k | minimal, low, medium, high | | `zai/glm-4.5-air` | GLM 4.5 Air | 128k | 96k | minimal, low, medium, high | | `zai/glm-4.5v` | GLM 4.5V | 66k | 16k | minimal, low, medium, high | | `zai/glm-4.6` | GLM 4.6 | 200k | 96k | minimal, low, medium, high | -| `zai/glm-4.6v` | GLM-4.6V | 128k | 24k | minimal, low, medium, high | -| `zai/glm-4.6v-flash` | GLM-4.6V-Flash | 128k | 24k | minimal, low, medium, high | | `zai/glm-4.7` | GLM 4.7 | 200k | 120k | minimal, low, medium, high | | `zai/glm-4.7-flash` | GLM 4.7 Flash | 200k | 131k | minimal, low, medium, high | | `zai/glm-4.7-flashx` | GLM 4.7 FlashX | 200k | 128k | minimal, low, medium, high | | `zai/glm-5` | GLM 5 | 202800 | 131100 | minimal, low, medium, high | | `zai/glm-5-turbo` | GLM 5 Turbo | 202800 | 131100 | minimal, low, medium, high | -| `zai/glm-5.1` | GLM 5.1 | 202k | 202k | minimal, low, medium, high | -| `zai/glm-5.2` | GLM 5.2 | 1040k | 128k | minimal, low, medium, high | +| `zai/glm-5.1` | GLM 5.1 | 202800 | 64k | minimal, low, medium, high | +| `zai/glm-5.2` | GLM 5.2 | 1000k | 128k | minimal, low, medium, high | | `zai/glm-5.2-fast` | GLM 5.2 Fast | 1000k | 128k | minimal, low, medium, high | +| `zai/glm-5.3` | GLM 5.3 | 1000k | 1000k | minimal, low, medium, high | +| `zai/glm-5.3-fast` | GLM 5.3 Fast | 1048576 | 262144 | minimal, low, medium, high | +| `zai/glm-5.3-flash` | GLM 5.3 Flash | 1000k | 131k | minimal, low, medium, high | +| `zai/glm-5.3-flashx` | GLM 5.3 FlashX | 1000k | 131072 | minimal, low, medium, high | | `zai/glm-5v-turbo` | GLM 5V Turbo | 200k | 128k | minimal, low, medium, high | ## xai | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `grok-4.3` | Grok 4.3 | 1000k | 30k | minimal, low, medium, high | +| `grok-4.3` | Grok 4.3 | 1000k | 30k | low, medium, high | | `grok-4.5` | Grok 4.5 | 500k | 500k | low, medium, high | -| `grok-build-0.1` | Grok Build 0.1 | 256k | 256k | minimal, low, medium, high | +| `grok-4.6` | Grok 4.6 | 500k | 500k | low, medium, high, xhigh | +| `grok-4.7` | Grok 4.7 | 500k | 500k | low, medium, high, xhigh | ## xiaomi-token-plan-ams | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `mimo-v2-pro` | MiMo-V2-Pro | 1048576 | 131072 | minimal, low, medium, high | | `mimo-v2.5` | MiMo-V2.5 | 1048576 | 131072 | minimal, low, medium, high | | `mimo-v2.5-pro` | MiMo-V2.5-Pro | 1048576 | 131072 | minimal, low, medium, high | +| `mimo-v2.6-flash` | MiMo-V2.6-Flash | 1048576 | 131072 | minimal, low, medium, high | +| `mimo-v2.6-pro` | MiMo-V2.6-Pro | 1048576 | 131072 | minimal, low, medium, high | ## xiaomi-token-plan-cn | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `mimo-v2-pro` | MiMo-V2-Pro | 1048576 | 131072 | minimal, low, medium, high | | `mimo-v2.5` | MiMo-V2.5 | 1048576 | 131072 | minimal, low, medium, high | | `mimo-v2.5-pro` | MiMo-V2.5-Pro | 1048576 | 131072 | minimal, low, medium, high | +| `mimo-v2.6-flash` | MiMo-V2.6-Flash | 1048576 | 131072 | minimal, low, medium, high | +| `mimo-v2.6-pro` | MiMo-V2.6-Pro | 1048576 | 131072 | minimal, low, medium, high | ## xiaomi-token-plan-sgp | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `mimo-v2-pro` | MiMo-V2-Pro | 1048576 | 131072 | minimal, low, medium, high | | `mimo-v2.5` | MiMo-V2.5 | 1048576 | 131072 | minimal, low, medium, high | | `mimo-v2.5-pro` | MiMo-V2.5-Pro | 1048576 | 131072 | minimal, low, medium, high | +| `mimo-v2.6-flash` | MiMo-V2.6-Flash | 1048576 | 131072 | minimal, low, medium, high | +| `mimo-v2.6-pro` | MiMo-V2.6-Pro | 1048576 | 131072 | minimal, low, medium, high | ## xiaomi | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `mimo-v2-flash` | MiMo-V2-Flash | 262144 | 65536 | minimal, low, medium, high | -| `mimo-v2-omni` | MiMo-V2-Omni | 262144 | 131072 | minimal, low, medium, high | -| `mimo-v2-pro` | MiMo-V2-Pro | 1048576 | 131072 | minimal, low, medium, high | | `mimo-v2.5` | MiMo-V2.5 | 1048576 | 131072 | minimal, low, medium, high | | `mimo-v2.5-pro` | MiMo-V2.5-Pro | 1048576 | 131072 | minimal, low, medium, high | | `mimo-v2.5-pro-ultraspeed` | MiMo-V2.5-Pro-UltraSpeed | 1048576 | 131072 | minimal, low, medium, high | +| `mimo-v2.6-flash` | MiMo-V2.6-Flash | 1048576 | 131072 | minimal, low, medium, high | +| `mimo-v2.6-pro` | MiMo-V2.6-Pro | 1048576 | 131072 | minimal, low, medium, high | +| `mimo-v2.6-pro-ultraspeed` | MiMo-V2.6-Pro-UltraSpeed | 1048576 | 131072 | minimal, low, medium, high | ## zai-coding-cn | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `glm-4.5-air` | GLM-4.5-Air | 131072 | 98304 | minimal, low, medium, high | -| `glm-4.7` | GLM-4.7 | 204800 | 131072 | minimal, low, medium, high | -| `glm-5-turbo` | GLM-5-Turbo | 200k | 131072 | minimal, low, medium, high | -| `glm-5.1` | GLM-5.1 | 200k | 131072 | minimal, low, medium, high | -| `glm-5.2` | GLM-5.2 | 1000k | 131072 | low, medium, high, max | -| `glm-5v-turbo` | GLM-5V-Turbo | 200k | 131072 | minimal, low, medium, high | +| `glm-4.6v` | GLM-4.6V | 128k | 32768 | minimal, low, medium, high | +| `glm-5.3` | GLM-5.3 | 1000k | 131072 | low, high, max | +| `glm-5.3-flash` | GLM-5.3-Flash | 1000k | 131072 | low, high, max | +| `glm-5.3-highspeed` | GLM-5.3 Highspeed | 1000k | 131072 | low, high, max | ## zai | Model | Name | Context | Max output | Reasoning levels | | --- | --- | --- | --- | --- | -| `glm-4.5-air` | GLM-4.5-Air | 131072 | 98304 | minimal, low, medium, high | | `glm-4.7` | GLM-4.7 | 204800 | 131072 | minimal, low, medium, high | | `glm-5-turbo` | GLM-5-Turbo | 200k | 131072 | minimal, low, medium, high | -| `glm-5.1` | GLM-5.1 | 200k | 131072 | minimal, low, medium, high | -| `glm-5.2` | GLM-5.2 | 1000k | 131072 | low, medium, high, max | -| `glm-5v-turbo` | GLM-5V-Turbo | 200k | 131072 | minimal, low, medium, high | +| `glm-5.2` | GLM-5.2 | 1000k | 131072 | high, max | +| `glm-5.2-highspeed` | GLM-5.2 Highspeed | 1000k | 131072 | high, max | +| `glm-5.3` | GLM-5.3 | 1000k | 131072 | low, high, max | +| `glm-5.3-flash` | GLM-5.3-Flash | 1000k | 131072 | low, high, max | +| `glm-5.3-highspeed` | GLM-5.3 Highspeed | 1000k | 131072 | low, high, max | diff --git a/scripts/write-models-md.mjs b/scripts/write-models-md.mjs index 45dd57a..ae4a992 100644 --- a/scripts/write-models-md.mjs +++ b/scripts/write-models-md.mjs @@ -6,6 +6,7 @@ import { writeFileSync } from "node:fs"; import path from "node:path"; import { fileURLToPath } from "node:url"; import { createPiModelRegistry } from "../dist/provider/provider-services.js"; +import { getAmbientCredentialNote, getPiApiKeyEnvVarNames } from "../dist/provider/pi-ai-models.js"; const stubAuthStorage = { loadAll: () => ({}), @@ -40,7 +41,16 @@ const lines = [ "The **Reasoning levels** column shows each model's native thinking levels from the registry (a dash means none).", "codegenie's `--reasoning` flag (and the `:reasoning` suffix) accepts `low`, `medium`, `high`, `xhigh`, or `auto`", "and maps onto whatever the model natively supports. Listing here means the model is known, not authenticated —", - "connect a provider with `codegenie provider login ` (or env vars / the Action's `llm-api-key`).", + "connect a provider with `codegenie provider login `, or set the env var listed under [Credentials](#credentials).", + "", + "## Credentials", + "", + "The env vars each provider's credentials are read from, in lookup order. In the GitHub Action, set them on the", + "`uses:` step's `env:`, or pass one key as `llm-api-key` when every configured model uses the same provider.", + "", + "| Provider | Credentials |", + "| --- | --- |", + ...[...byProvider.keys()].map((provider) => `| \`${provider}\` | ${credentialsCell(provider)} |`), "" ]; @@ -55,6 +65,13 @@ for (const [provider, group] of byProvider) { lines.push(""); } +function credentialsCell(provider) { + const envVars = getPiApiKeyEnvVarNames(provider).map((name) => `\`${name}\``).join(", "); + const note = getAmbientCredentialNote(provider); + const cell = [envVars, note].filter((part) => part !== undefined && part !== "").join(" "); + return escapeCell(cell === "" ? "—" : cell); +} + function formatTokens(value) { if (value === undefined || value <= 0) { return "—"; diff --git a/specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md b/specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md new file mode 100644 index 0000000..59c0f17 --- /dev/null +++ b/specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md @@ -0,0 +1,271 @@ +# Issue 125: GitHub Action Model Aliases and Provider Credential Docs + +Status: IMPLEMENTED (dogfood pending) +Based on: owner request (2026-09-25) for a set of named models in the Action, selectable per comment trigger +Depends on: plan 97 (GitHub Action adapter, trust model, status comment); `src/provider/pi-ai-models.ts` env-key routing + +## Objective + +Let a workflow author configure several named models once. The configured default runs automatic reviews and bare trigger comments. A collaborator can pick another configured model for one run by writing ` ` (e.g. `codegenie review opus`). At the same time, document every provider credential env var from the table codegenie actually reads, and fix the drift in that table. + +Workflows that use only `model` + `llm-api-key` keep working, with one deliberate change: **`llm-api-key` now takes precedence**. If it is set, it is the only model key used and it overrides provider env vars. If it is not set, provider credentials resolve exactly as today. + +Target surface, with several providers (per-provider env vars, no `llm-api-key`): + +```yaml +- uses: 0xPolygon/codegenie@ + env: + # Credential env var per provider: see models.md#credentials + OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }} + ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} + with: + model: deepseek # default: automatic reviews + bare "codegenie review" + models: | + deepseek: openrouter/deepseek/deepseek-v4.1-flash:max + astra: openrouter/openai/gpt-6-astra:medium + glm: openrouter/z-ai/glm-5.3:max + opus: anthropic/claude-opus-5 +``` + +With a single provider, one `llm-api-key` covers every alias: + +```yaml + with: + model: deepseek + models: | + deepseek: openrouter/deepseek/deepseek-v4.1-flash:max + glm: openrouter/z-ai/glm-5.3:max + llm-api-key: ${{ secrets.OPENROUTER_API_KEY }} +``` + +The simple form is unchanged: `model: "openrouter/deepseek/deepseek-v4.1-flash:max"` + `llm-api-key: ${{ secrets.LLM_API_KEY }}`. + +## Evidence and constraints + +1. **Action inputs are strings.** A mapping under `with:` fails workflow validation, so `models` must be a block string that codegenie parses. `models` holds model specs only; keys never go in it. +2. **Trailing comment text is ignored today.** `matchesTriggerPhrase` accepts `` followed by whitespace and discards the rest (`src/github-action/event-gate.ts`). Plan 97 (line 88) deferred comment-driven options to "an allowlisted grammar if ever wanted". This plan is that grammar: a closed-set lookup into aliases defined by the workflow author. Comment text never becomes a model spec, reasoning level, or flag. +3. **`llm-api-key` currently loses to env vars, and follows whichever model is selected.** `applyGenericApiKey` copies `LLM_API_KEY` into the selected provider's env var only when that var is empty, so a stale `ANTHROPIC_API_KEY` in the job env silently beats the key the workflow passed explicitly. The owner wants the reverse: an explicitly passed `llm-api-key` is authoritative. Separately, with aliases, one key would be copied into whatever provider the selected alias uses. An OpenRouter key would land in `ANTHROPIC_API_KEY` for `codegenie review opus` and fail with a confusing provider 401. +4. **The missing-credential error gives CLI advice in CI.** `adapter.resolveModel` (`resolveRealModel`) returns the same `undefined` for an unknown model, a deprecated model, a provider with no usable credentials, and any swallowed exception. `createPiRunner` then throws `no usable LLM model could be resolved; run codegenie provider login …` (`src/llm/pi-runner.ts:220`). The status comment's failure body shows only the error code and a provider message (`renderFailureBody`), so a CI user sees `config_error` and nothing actionable. +5. **The provider→env-var table has drifted.** `API_KEY_ENV_VARS` in `src/provider/pi-ai-models.ts` is a hand copy of pi-ai's table (pi-ai 0.87.1 does not export `env-api-keys`). It lacks `qwen-token-plan` / `qwen-token-plan-cn` / `qwen-token-plan-individual` (`QWEN_TOKEN_PLAN_API_KEY`, `QWEN_TOKEN_PLAN_CN_API_KEY`), `baseten`, `meta` and `radius`. The two qwen providers are listed in `models.md` today. Their env keys are never read, and `llm-api-key` fails for them with "provider does not accept an API key". No test covers the table. +6. **Credential names are undocumented and non-uniform.** + - Unusual names: `google` → `GEMINI_API_KEY`, `vercel-ai-gateway` → `AI_GATEWAY_API_KEY`, `huggingface` → `HF_TOKEN`, `github-copilot` → `COPILOT_GITHUB_TOKEN`. + - Shared names: `moonshotai*` → `MOONSHOT_API_KEY`, `opencode*` → `OPENCODE_API_KEY`, `cloudflare-*` → `CLOUDFLARE_API_KEY`. + - Other credentials: `amazon-bedrock` uses AWS credentials or OIDC; `google-vertex` uses `GOOGLE_CLOUD_API_KEY`, or ADC + `GOOGLE_CLOUD_PROJECT` + `GOOGLE_CLOUD_LOCATION`. + - `openai-codex` has only a stored OAuth login, which works on a runner only where someone ran `codegenie provider login` (self-hosted). +7. **The examples need two copies of the list.** The automatic and comment examples are separate workflow files. The dogfood workflow (`.github/workflows/codegenie-review.yml`) already shows that one workflow can serve both events without changing the trust model. `ref: ${{ github.event.pull_request.base.sha || '' }}` checks out the base SHA on `pull_request` and the default branch on `issue_comment`, and the concurrency group is keyed by `pull_request.number || issue.number`. + +## 1. Model aliases in the Action + +### Configuration + +- New input `models`: a block string holding a flat YAML mapping of `alias: provider/model[:reasoning]`. Parse it with the existing `yaml` dependency. Reject anything other than string keys mapped to string values (no nesting, anchors or lists). Blank input means no aliases. +- Alias names accept ASCII letters in either case: `^[A-Za-z0-9][A-Za-z0-9._-]{0,31}$`. Normalize to lowercase before duplicate checks, storage and lookup, including the default `model` alias. `Opus` and `opus` are duplicates; a comment requesting `OPUS` selects `opus`. +- Each value must pass `parseModelSpec` **with** a provider prefix, and the provider must exist in the registry. The whole block fails if any entry is invalid, since that is an error in the workflow file. Such a config error fails every run, including automatic reviews, with an Actions log message naming the line and alias. +- Preserve `splitReasoningSuffix` semantics: + - Only a recognized reasoning suffix (including `auto`) is separated from the model id. + - Other suffixes, such as `:free` or `:8b`, remain part of the id. + - A typo such as `:hgh` therefore surfaces as an unknown model when that alias is selected (see Credentials), not as an alias-parse error. + - Model-specific reasoning handling stays in the existing review path. +- `model` stays the single source of the default. It accepts either an alias name or a full spec. If `models` is set: + - `model` is required. + - A `model` value without a `/` must name an alias; otherwise it is a config error, because it is almost certainly a typo. + - Without `models`, `model` parsing is unchanged. +- There is no reserved `default` alias and no per-alias setting beyond the spec. + +### Trigger grammar + +- `event-gate.ts` stays pure. On a phrase match it also returns `requestedAlias`: the first whitespace-delimited token on the **first line** after the phrase, lowercased, or none. Later lines and later tokens are ignored, so `codegenie review\n\nfocus on auth` still means "default". +- Resolution happens in the entrypoint, **after** the payload authorization and the live write-permission check, so unauthorized commenters never get a reply: + +| `models` set? | Comment | Result | +| --- | --- | --- | +| no | any match | Today's behavior: `model`, trailing text ignored | +| yes | `` | default (`model`) | +| yes | ` ` | that alias | +| yes | ` ` | reply (below), no review | + +- The `pull_request` lane always uses the default. +- Comments never accept raw specs, `:reasoning` suffixes or extra flags. Authors who want a reasoning variant define an alias for it (`opus-max: anthropic/claude-opus-5:max`). + +### Unknown alias reply + +- Post one **new** plain comment with `createComment`. The status comment is never touched, so the last report stays intact, and the reply carries no status marker, so it can never be reclaimed. +- Fixed text built only from the configured alias names, e.g. ``Unknown model. Available: `deepseek` (default), `astra`, `glm`, `opus`.`` The commenter's token is not echoed. Alias names are restricted to `[A-Za-z0-9._-]`, so the body needs no sanitizing beyond the normal posting path. +- The run ends as a skip (exit 0) with decision reason `unknown model alias`. + +### Preflight + +- `preflight-only` authorizes the event and resolves the alias. For an unknown alias it posts the reply and sets `should-run=false`, so the reply is posted once. It needs no LLM credentials. +- Credential handling (below) happens only in the review invocation. A two-job workflow must pass the same `model` / `models` / trigger inputs to both jobs; the review job repeats authorization and resolution. + +### Credentials + +There are two modes, chosen by whether `llm-api-key` is set (the input, or `LLM_API_KEY` in the step env, as today). + +**`llm-api-key` set: it is the only key.** + +- Every selectable model (the default and every alias) must share one provider, and that provider must accept an API key. Check the whole set regardless of which alias was selected, and fail with a config error otherwise: + + > llm-api-key is a single key, but models use providers openrouter, anthropic. Remove llm-api-key and set each provider's env var (see models.md#credentials). + + A key that works for some aliases and 401s for others is worse than failing up front. +- Routing: + - Unset that provider's other credential env vars: every name in its table entry, plus `ANTHROPIC_AUTH_TOKEN` for Anthropic. pi's Anthropic auth checks `ANTHROPIC_AUTH_TOKEN` first and sends it as a Bearer header. + - Then write the key into the provider's API-key env var unconditionally. + - Other providers' vars are left alone. + - Amendment (implementation review, owner decision 2026-09-25): the original claim here was wrong. codegenie's `resolveProviderAuth` checks env first, but in production pi-ai's own auth resolution (`dist/auth/resolve.js`) lets a stored credential own its provider ahead of env vars, because `complete()` passes no explicit `apiKey`. So when `llm-api-key` is routed and the runner has a stored codegenie login for that provider (self-hosted runners only), the Action fails before claiming the status comment, telling the user to run `codegenie provider logout ` or unset `llm-api-key`. The Action-only guard was chosen over passing the key to pi explicitly, which would also flip CLI precedence (env over stored login) and could silently change CLI billing. +- Providers without an API key (`amazon-bedrock`, `openai-codex`) keep today's error ("does not accept an API key"). + +**`llm-api-key` not set:** resolution is unchanged (env vars, then ambient credentials, then any stored login), with no Action-specific pre-check and no new policy. Aliases that are not selected are never touched, so a missing Anthropic key never breaks an automatic review whose default uses another provider. + +**A clear error when credentials are missing.** This one fix serves both the CLI and CI: + +- Keep the failure reason where resolution happens. `resolveRealModel` (`src/llm/pi-runner.ts`) currently returns `undefined` for four different failures: a model missing from the registry, a deprecated model (`isDeprecatedProviderModel`), a provider without credentials, and any exception swallowed by its `catch {}`. Have it report which one occurred (a small reason alongside the `undefined` result; the resolution logic itself is unchanged), so `createPiRunner` can pick the message: + - unknown model: "unknown model `provider/id`"; + - deprecated model: "model `provider/id` is deprecated"; + - missing credentials: "no credentials for provider `anthropic`: set `ANTHROPIC_API_KEY` (or run `codegenie provider login anthropic`)"; + - unexpected error: today's generic message, unchanged. Never label it missing credentials. + + All four keep the existing `config_error` code; only the message text changes. +- Take the env var name(s) from the phase 2 table, and use the ambient note for Bedrock, Vertex and Codex. +- This fires after the status comment is claimed. It reaches the PR through the existing failure path, the same way any failed run replaces the status comment today. +- `renderFailureBody` renders the message only for these model-resolution errors, not for codegenie errors in general, which can carry git stderr, paths or other external text. The message can still include workflow-supplied model ids, so pass it through `scrubGitHubSecrets`. +- Sanitize the whole failure body. Today `finalizeFailure` sends `renderFailureBody` output to GitHub without `sanitizeGitHubCommentBody` (only the success report path is sanitized, `src/github-action/status-comment.ts`), so an existing provider message can carry mentions or HTML. Wrapping the rendered failure body in the existing sanitizer closes that gap and covers the new message. No new reply path and no new helpers. + +**Compatibility.** + +- Single-model workflows that set only `llm-api-key`, or only native env vars, behave as before. +- When both are set with different values for the same provider, `llm-api-key` now wins. Document this in the README and release notes, and update the `action.yml` `llm-api-key` description. +- The `inputs.llm-api-key || env.LLM_API_KEY` fallback in `action.yml` stays. + +No other new key input. A provider-keyed `llm-api-keys` input is a Future Consideration, only if env var names prove to be a real stumbling block after phase 2's docs. + +### Wiring + +- `action.yml`: + - Add the `models` input and pass it as `--models "$INPUT_MODELS"` (no secrets, so argv is fine). + - Update the `model` description: alias or spec. + - Update the `llm-api-key` description: when set, it is the only model key and overrides provider env vars; all selectable models must share its provider. +- `parseGitHubActionArgs`: accept `--models`, and validate the combination with `model` after all flags are parsed. +- Add `modelAlias` and the resolved `modelSpec` to the decision and lifecycle records. The report already renders `🤖 Model:`, and the trigger comment shows who asked, so there is no extra rendered line. +- No new flags on `codegenie review`. The synthesized argv is unchanged in shape: `--provider/--model/--reasoning` from the resolved spec. + +### Tests and acceptance + +- `models` parsing: + - Valid blocks, including quoted values, comments and blank lines. + - Mixed-case aliases normalize consistently; case-insensitive duplicates, bad names, nested or list values are rejected. + - Missing provider prefix and unknown provider are rejected. Recognized reasoning suffixes parse normally; colon-bearing ids stay intact: `…:free` parses as model id ending in `:free` with the default reasoning, and `…:free:high` as model id ending in `:free` with reasoning `high`. + - YAML syntax errors. +- `model` resolution: an alias name, a full spec, a slash-less `model` that is not an alias while `models` is set, and a missing `model` while `models` is set. +- Gate: the token extracted from the first line only; multi-line comments; a phrase inside a longer word (still no match); unchanged behavior when `models` is absent. +- Entrypoint: + - The PR lane uses the default. + - A bare comment uses the default; an alias comment produces the alias's argv. + - An unknown alias posts exactly one reply, with no status claim, no review and exit 0. + - An unauthorized actor with an unknown alias is skipped silently. + - `preflight-only`: an unknown alias gives `should-run=false` plus one reply; a known alias passes without LLM credentials. +- Credentials: + - With `llm-api-key` set, it overrides a different pre-set native var for the same provider, and the competing vars are unset. The resolved key equals `llm-api-key` even with a stored credential present. Tests isolate `process.env`, because `providerEnvValue` falls back to it. + - Models spanning several providers with `llm-api-key` set fail with the config error. + - A single-provider alias set works for every alias. + - A keyless provider errors as before. + - With `llm-api-key` unset, each alias uses its own provider's native var. + - Resolution failures are labeled correctly: + - an unknown model id gives the unknown-model message; + - a deprecated model gives the deprecated message; + - a known model without credentials gives the provider + env-var message in the status comment's failure body; + - an unexpected resolver exception gives the generic message, not "missing credentials". + - Failure-body sanitizing: a failure message containing an `@mention`, HTML comment and a secret-shaped string renders with the mention neutralized, the HTML stripped and the secret scrubbed. This covers both the new message and an existing provider message. +- Compatibility regression: the README's simple form (`model: openrouter/deepseek/deepseek-v4.1-flash:max`, `llm-api-key` set, no `models`) produces the same argv, the same `OPENROUTER_API_KEY`, and the same ignored trailing comment text as before. Also cover the native-env-only form (`OPENROUTER_API_KEY` in step env, no `llm-api-key`). +- Workflow shape: the merged example passes `actionlint`. + +Acceptance: + +- An existing single-model workflow produces identical argv and behavior, except that `llm-api-key` now wins over a conflicting native var for the same provider. +- With aliases, every authorized matching comment either runs a model the workflow author listed or gets the fixed reply. Unauthorized commenters get no reply. + +Likely files: `src/github-action/event-gate.ts`, `src/github-action/entrypoint.ts`, a small `src/github-action/models.ts` (parse and resolve), `src/github-action/render.ts`, `src/llm/pi-runner.ts` (resolution reason and messages), `src/github-action/status-comment.ts` (failure-body sanitizing), `action.yml`, `tests/github-action.test.ts`. + +## 2. Provider credential table: fix, guard, document + +### Fix and guard + +- Sync `API_KEY_ENV_VARS` with pi-ai 0.87.1: add `qwen-token-plan`, `qwen-token-plan-cn`, `qwen-token-plan-individual`, `baseten`, `meta` and `radius`. +- Keep the existing, deliberate exclusion of `ANTHROPIC_AUTH_TOKEN` from the lookup: pi's `getEnvApiKey` also skips it because it is bearer auth, and `getPiEnvApiKey` returns the first var it finds. Phase 1's routing unsets it separately. +- Add a small ambient-credential note table in the same module for `amazon-bedrock`, the ADC path of `google-vertex`, and `openai-codex` (stored login; self-hosted runners only). Export only what the generator, the phase 1 error and the coverage test need. +- **Registry coverage test:** every provider from `createPiModelRegistry(...).listProviders()` has env var names or an ambient note. This catches the drift actually observed: a provider added upstream with no key mapping. +- Future Consideration: ask pi-ai upstream to export its env-key table, then delete our copy. Deletion is the success outcome. + +### Document + +- `scripts/write-models-md.mjs`: + - Emit one `## Credentials` table near the top of `models.md`: provider → env var(s), or its ambient note. + - Replace the header sentence that mentions "env vars / the Action's `llm-api-key`" with a link to it. + - Regenerate `models.md`. +- README, GitHub Action section: + - Show the multi-model example above plus the unchanged simple form. + - Add a short credentials table for the common providers (anthropic, openai, openrouter, google → `GEMINI_API_KEY`, bedrock, vertex), flag the names that break the pattern, and link to `models.md#credentials`. + - Document: + - the alias grammar and the unknown-alias reply; + - the two credential modes: `llm-api-key` is the only key when set and needs a single-provider model set; otherwise provider credentials resolve as usual; + - the precedence change; + - concurrency and the status comment: a `codegenie review opus` comment supersedes an in-flight run on the same PR (newest event wins), and the single status comment shows the newest report. + - **Note: a typo'd alias still cancels an in-flight review.** `codegenie review opsu` starts a run that supersedes the running one, then only posts the reply. This is the same "any comment supersedes" residual plan 97 accepted. + - **Note: inline comments accumulate across models.** A `review opus` run after an automatic review posts its own inline findings. Duplicate suppression only matches unchanged prose, so two models reporting the same issue both appear. This is expected for a second opinion. +- Examples: + - Replace `codegenie-review-pr.yml` and `codegenie-review-comment.yml` with one `examples/workflows/codegenie-review.yml` that has both triggers, using the dogfood workflow's `ref` and concurrency expressions. + - Keep the trust-model and cancellation comments. + - Show the `models` block and native env keys, with a `# see models.md#credentials` pointer. + - Keep the simple single-model form as a commented alternative. + - Update README references to "both trigger lanes". +- Dogfood workflow: works unchanged. Adding aliases there needs repository secrets and is the owner's call, not part of this implementation. + +### Tests and acceptance + +- The coverage test fails on the pre-fix table and passes after the sync. +- `models.md` regenerates deterministically and its Credentials table lists every provider. + +## Implementation and validation order + +1. Phase 2 fix and guard first: it is a standalone bug fix, and phase 1's missing-credential message needs the env var names. +2. Phase 1: parsing, gate token and resolution, the reply, `llm-api-key` routing and the error split, then wiring and records. +3. Generator, `models.md` regeneration, README and the merged example. +4. Run `pnpm test`, `pnpm run typecheck`, `pnpm run build`, `make models-list` and `git diff --check`. Review for: + - comment text reaching argv beyond a listed alias; + - replies to unauthorized actors; + - `llm-api-key` precedence (it wins, and no competing credential for that provider remains readable); + - model-id suffix compatibility and case-insensitive alias resolution; + - any change to single-model behavior. +5. Dogfood on a test PR, run by the owner. Try a bare trigger, an alias trigger, an unknown alias, and an alias whose provider credentials are missing. Check the reply text, the status comment, and the decision/lifecycle records. Release follows the plan 97 order (npm before tag). + +## Scope and non-goals + +- No per-alias access control; the author's list is the allowlist and the cost control. +- No comment-supplied specs, reasoning levels or flags. +- No per-alias status comments, and no `llm-api-keys` input in this iteration. +- No Action-specific credential pre-check or stored-login policy, beyond the narrow `llm-api-key` stored-login guard (see the Credentials amendment). +- No aliases in `codegenie.toml` or the CLI. +- No review-pipeline, prompt or model-facing changes. +- Harness touches stay within the plan 97 seams, plus `pi-ai-models.ts` exports (existing seam), the `resolveRealModel` failure reason and `createPiRunner` messages, and the docs generator. + +## Separate follow-up (not this plan) + +- Unverified and older than this plan: with a custom GitHub App token, creating the status comment on the first run is itself an `issue_comment` event. GitHub starts workflows for App-token events, not for `GITHUB_TOKEN`. The new run skips as a bot, but under `cancel-in-progress: true` it may cancel the review that created the comment. Verify with an App-token setup; fix under plan 97 if confirmed. + +## Implementation and review results + +- **Aliases:** `src/github-action/models.ts` parses `models` with the YAML failsafe schema (every scalar is a string, so `405:` stays a name), resolves `model` as an alias or spec, selects the comment-requested alias, and renders the fixed unknown-alias reply. The event gate returns the first token on the trigger phrase's line. The entrypoint resolves it only after the live permission check. Decision and lifecycle records carry `modelAlias`/`modelSpec`. +- **`llm-api-key`:** it clears the provider's credential vars (including `ANTHROPIC_AUTH_TOKEN`), writes the key unconditionally, requires one API-key provider across all selectable models, and refuses when the runner has a stored login for that provider (amendment above). +- **Resolution errors:** `resolveRealModel` keeps its reason. The real adapter exposes it through an optional `explainUnresolvedModel`, so test adapters are unchanged. `createPiRunner` reports unknown model, deprecated model, missing credentials (naming the env var) or the unchanged generic message, all as `config_error` with `context.modelResolution`. The Action publishes only those messages, scrubbed and capped. The failure JSON gets a separate `modelResolution` field. The whole failure comment is now sanitized. +- **Credentials table:** synced with pi-ai 0.87.1, plus ambient notes and a registry coverage test. `models.md` gained a generated Credentials table. Regenerating it also caught up with the installed registry: 1103 → 1490 models, 37 → 41 providers. The table sync matters because `baseten`, `meta`, `radius` and `qwen-token-plan-individual` are live providers. +- **Docs and examples:** README, `action.yml` and one merged example workflow. The owner's in-flight switch of the example default to `openrouter/openai/gpt-6-luna:xhigh` was carried into the merged example as the `luna` alias. +- **Independent review (fresh agent, adversarial):** + - Found that pi-ai's stored-credential-first auth bypasses `llm-api-key` on self-hosted runners with a stored login. Resolved by the owner-chosen Action guard. + - Found that a nonexistent provider was reported as "missing credentials". Fixed: it now gets the generic message. + - Found that numeric alias names were rejected. Fixed with the failsafe schema. + - Not changed: the example and README pins stay `@v0.6.3`, which has no `models` input. The existing contract test forces them to the package version at release. The optional adapter method (rather than a reason-returning `resolveModel`) was kept to avoid touching the ~40 test adapters. + - Sound: the trust boundary, secrets on every published surface, backward compatibility and error labeling on the Action path. +- **Mutation checks:** removing the failure-body sanitizer, or the credential clearing, fails the new tests. + +Validation: `pnpm test` passed **1,583 tests across 69 files**, including actionlint on the merged example. `make evals` passed **39 synthetic tests**. Typecheck, build and `git diff --check` passed, and `models.md` regenerates identically. The built CLI printed the missing-credentials, unknown-model and deprecated-model messages on a real commit range, failing before any model call. No paid inference, pushes or GitHub posts were made. Dogfooding on a test PR (bare trigger, alias, unknown alias, alias with missing credentials) and the release remain with the owner. diff --git a/specs/plans/README.md b/specs/plans/README.md index 5db05e5..7b2f189 100644 --- a/specs/plans/README.md +++ b/specs/plans/README.md @@ -127,6 +127,7 @@ This directory tracks implementation plans for confirmed improvements. Status va | 122 | CORE IMPLEMENTED (live baseline pending; C2 deferred) | [Issue 122: Shared Evidence and Focused Review Follow-ups](122-issue-122-shared-evidence-and-focused-review-followups.md) | | 123 | IMPLEMENTED (live comparisons pending) | [Issue 123: Reliable Search Evidence and Honest Review Outcomes](123-issue-123-search-evidence-and-honest-review-outcomes.md) | | 124 | IMPLEMENTED | [Issue 124: Repair Feedback, Evidence Retention, and Faithful Fallback Reports](124-issue-124-repair-feedback-evidence-retention-and-fallback-reports.md) | +| 125 | IMPLEMENTED (dogfood pending) | [Issue 125: GitHub Action Model Aliases and Provider Credential Docs](125-issue-125-github-model-aliases-and-provider-credentials.md) | ## Recommended order for 106-110 diff --git a/src/github-action/entrypoint.ts b/src/github-action/entrypoint.ts index 47e3274..c709568 100644 --- a/src/github-action/entrypoint.ts +++ b/src/github-action/entrypoint.ts @@ -1,7 +1,8 @@ import { appendFileSync, readFileSync, writeFileSync } from "node:fs"; import path from "node:path"; import { executeReviewCommand, parseReviewCommand } from "../cli/review-command.js"; -import { getPiApiKeyEnvVarName } from "../provider/pi-ai-models.js"; +import { getCodegeniePaths } from "../config/paths.js"; +import { createFileAuthStorage, type PiAuthStorage } from "../provider/provider-services.js"; import { renderMarkdownReview } from "../output/markdown-renderer.js"; import { sanitizeGitHubCommentBody, scrubGitHubSecrets } from "../github/comment-sanitizer.js"; import type { ReviewResult, TelemetryEvent } from "../types.js"; @@ -17,8 +18,17 @@ import { type TriggerDecision, type TriggerRules } from "./event-gate.js"; -import { splitReasoningSuffix } from "../provider/reasoning.js"; import { createIssueCommentClient, type IssueCommentClient } from "./issue-comments.js"; +import { + applyLlmApiKey, + formatModelSpec, + parseModelAliases, + renderUnknownAliasReply, + resolveModelConfig, + selectModel, + type ModelConfig, + type ModelSelection +} from "./models.js"; import { createStatusCommentController } from "./status-comment.js"; import { renderProviderMessage, renderStructuredSubmitFailure } from "./render.js"; @@ -49,6 +59,8 @@ export type ExecuteGitHubActionOptions = { issueComments?: IssueCommentClient; runReview?: (reviewArgv: string[], hooks: ReviewHooks) => Promise; minEditIntervalMs?: number; + // Stored codegenie logins on this runner (self-hosted); tests inject one. + authStorage?: Pick; }; type GitHubActionInputs = { @@ -59,7 +71,7 @@ type GitHubActionInputs = { postInlineComments: boolean; preflightOnly: boolean; botLogin?: string; - model?: ModelSpec; + models: ModelConfig; reviewPassthrough: string[]; }; @@ -69,11 +81,7 @@ type AuthorizedDecision = Extract & { permissionCheck: PermissionCheck; }; -export type ModelSpec = { - provider?: string; - model: string; - reasoning: string; -}; +export { applyLlmApiKey, parseModelSpec, type ModelSpec } from "./models.js"; // The `codegenie github-action` subcommand: the whole GitHub Actions surface // (plan 97). Composes the review path through its public seams only — the @@ -135,23 +143,47 @@ export async function executeGitHubActionCommand( } const authorized: AuthorizedDecision = { ...decision, permissionCheck }; - writePreflightOutputs(env, true, decision.prNumber); - writeDecisionRecord(write, { + const decisionFields = { eventName, - run: true, lane: decision.lane, prNumber: decision.prNumber, actor: decision.actor, association: decision.association, actorAllowlisted: decision.actorAllowlisted, permissionCheck - }); + }; + + // Resolved only after authorization, so unauthorized commenters never get + // a reply. An unlisted alias gets fixed text from the workflow's own list. + const selected = selectModel(inputs.models, decision.requestedAlias); + if (selected.kind === "unknown_alias") { + await comments.createComment(decision.prNumber, renderUnknownAliasReply(inputs.models)); + const reason = "unknown model alias"; + write(`github-action: skipped — ${reason}\n`); + writeDecisionRecord(write, { ...decisionFields, run: false, reason }); + writePreflightOutputs(env, false); + return; + } + const selection = selected.selection; + writePreflightOutputs(env, true, decision.prNumber); + writeDecisionRecord(write, { ...decisionFields, run: true, ...modelRecordFields(selection) }); if (inputs.preflightOnly) { write(`github-action: preflight authorized ${decision.lane} trigger for PR #${decision.prNumber}\n`); return; } - applyGenericApiKey(env, inputs.model); + const keyProvider = applyLlmApiKey(env, inputs.models); + if (keyProvider !== undefined) { + // pi-ai lets a stored login own its provider ahead of env vars, which + // would silently bypass llm-api-key. Only self-hosted runners can have one. + const storage = opts.authStorage ?? createFileAuthStorage(getCodegeniePaths(undefined, env)); + if (storage.get(keyProvider) !== undefined) { + throw new CodegenieError( + "invalid_args", + `a stored codegenie login for ${keyProvider} on this runner would override llm-api-key; run \`codegenie provider logout ${keyProvider}\` on the runner, or unset llm-api-key to use the stored login` + ); + } + } // Identity resolution order: explicit bot-login input (custom GitHub // Apps) → /user lookup (PATs) → the GITHUB_TOKEN default. Reclaim and @@ -181,13 +213,13 @@ export async function executeGitHubActionCommand( String(decision.prNumber), "--ci", ...(inputs.postInlineComments ? ["--post-github-comments"] : []), - ...(inputs.model !== undefined + ...(selection !== undefined ? [ - ...(inputs.model.provider !== undefined ? ["--provider", inputs.model.provider] : []), + ...(selection.spec.provider !== undefined ? ["--provider", selection.spec.provider] : []), "--model", - inputs.model.model, + selection.spec.model, "--reasoning", - inputs.model.reasoning + selection.spec.reasoning ] : []), ...inputs.reviewPassthrough @@ -210,19 +242,24 @@ export async function executeGitHubActionCommand( const code = actionErrorCode(error); const diagnostic = structuredSubmitFailureDiagnosticFromError(error); const providerMessage = providerMessageFromError(error); + const modelResolution = modelResolutionFromError(error); + // Our own model-resolution text rides the same explanation slot as a + // provider message; the failure JSON keeps the two distinct. + const explanation = providerMessage ?? modelResolution?.message; publishFailureFiles({ errorCode: code, decision: authorized, env, ...(diagnostic !== undefined ? { diagnostic } : {}), ...(providerMessage !== undefined ? { providerMessage } : {}), + ...(modelResolution !== undefined ? { modelResolution } : {}), ...(runUrl !== undefined ? { runUrl } : {}) }); - await controller.finalizeFailure(code, diagnostic, providerMessage); - emitActionRecord(attachment?.runDir, eventName, authorized, "review_failed", controller.stats(), env, write, code); + await controller.finalizeFailure(code, diagnostic, explanation); + emitActionRecord(attachment?.runDir, eventName, authorized, selection, "review_failed", controller.stats(), env, write, code); const detail = diagnostic !== undefined ? renderStructuredSubmitFailure(diagnostic) : code; write( - `github-action: review failed — ${detail}${providerMessage !== undefined ? `: ${providerMessage}` : ""}\n` + `github-action: review failed — ${detail}${explanation !== undefined ? `: ${explanation}` : ""}\n` ); throw error; } @@ -232,17 +269,17 @@ export async function executeGitHubActionCommand( publishReportFiles(runResult.reportMarkdown, env); if (runResult.failed) { await controller.finalizeFailure("review_failed", undefined, "Required review work failed. See the saved report for diagnostics.", runResult.reportMarkdown); - emitActionRecord(runResult.runDir, eventName, authorized, "review_failed", controller.stats(), env, write, "review_failed"); + emitActionRecord(runResult.runDir, eventName, authorized, selection, "review_failed", controller.stats(), env, write, "review_failed"); throw new CodegenieError("review_failed", "Required review work failed; partial report retained."); } try { await controller.finalizeSuccess(runResult.reportMarkdown); } catch (error) { const code = actionErrorCode(error); - emitActionRecord(runResult.runDir, eventName, authorized, "terminal_post_failed", controller.stats(), env, write, code); + emitActionRecord(runResult.runDir, eventName, authorized, selection, "terminal_post_failed", controller.stats(), env, write, code); throw error; } - emitActionRecord(runResult.runDir, eventName, authorized, "success", controller.stats(), env, write); + emitActionRecord(runResult.runDir, eventName, authorized, selection, "success", controller.stats(), env, write); write(`github-action: review complete — report posted to PR #${decision.prNumber}\n`); } @@ -275,8 +312,11 @@ export function parseGitHubActionArgs(argv: string[]): GitHubActionInputs { allowedUsers: [], postInlineComments: true, preflightOnly: false, + models: { aliases: new Map() }, reviewPassthrough: [] }; + let modelInput: string | undefined; + let modelsInput = ""; const passthroughFlags = new Set(["--depth", "--lens", "--max-time", "--budget-boost"]); for (let index = 0; index < argv.length; index += 2) { @@ -298,9 +338,9 @@ export function parseGitHubActionArgs(argv: string[]): GitHubActionInputs { } else if (flag === "--preflight-only") { inputs.preflightOnly = parseBoolean(flag, value); } else if (flag === "--model") { - if (value.trim() !== "") { - inputs.model = parseModelSpec(value.trim()); - } + modelInput = value; + } else if (flag === "--models") { + modelsInput = value; } else if (flag === "--bot-login") { if (value.trim() !== "") { inputs.botLogin = value.trim(); @@ -316,55 +356,11 @@ export function parseGitHubActionArgs(argv: string[]): GitHubActionInputs { if (inputs.triggerPhrase.trim() === "") { throw new CodegenieError("invalid_args", "--trigger-phrase must not be empty"); } + // `model` and `models` are validated together, after every flag is read. + inputs.models = resolveModelConfig(modelInput, parseModelAliases(modelsInput)); return inputs; } -// One model spec instead of separate provider/model/reasoning inputs: -// `provider/model[:reasoning]`, e.g. `anthropic/claude-opus-5:xhigh`. -// Reasoning defaults to "high" — Action reviews are unattended, so the -// action's posture favors quality over the CLI's interactive default. The -// suffix rule is the shared one (a non-level `:suffix` stays in the model id). -export function parseModelSpec(spec: string): ModelSpec { - const { model: rest, reasoning = "high" } = splitReasoningSuffix(spec); - const slash = rest.indexOf("/"); - const provider = slash > 0 ? rest.slice(0, slash) : undefined; - const model = slash > 0 ? rest.slice(slash + 1) : rest; - if (model.trim() === "" || (slash === 0)) { - throw new CodegenieError("invalid_args", `--model must be provider/model[:reasoning], got: ${spec}`); - } - return { - ...(provider !== undefined ? { provider } : {}), - model, - reasoning - }; -} - -// Routes a generic LLM_API_KEY to the env var the selected provider actually -// reads (the names are not uniform: google reads GEMINI_API_KEY). Explicitly -// set provider-native vars always win; the generic key never overwrites. -export function applyGenericApiKey(env: NodeJS.ProcessEnv, model: ModelSpec | undefined): void { - const generic = env.LLM_API_KEY; - if (generic === undefined || generic === "") { - return; - } - if (model?.provider === undefined) { - throw new CodegenieError( - "invalid_args", - "LLM_API_KEY requires a model input with a provider prefix (e.g. anthropic/claude-opus-5) so the key can be routed" - ); - } - const envVarName = getPiApiKeyEnvVarName(model.provider); - if (envVarName === undefined) { - throw new CodegenieError( - "invalid_args", - `provider ${model.provider} does not accept an API key; set its native credentials instead of LLM_API_KEY` - ); - } - if (env[envVarName] === undefined || env[envVarName] === "") { - env[envVarName] = generic; - } -} - async function hasWritePermission(comments: IssueCommentClient, login: string): Promise { if (login === "") { return false; @@ -435,8 +431,11 @@ type ActionFailureRecord = { runId?: string; structuredSubmitFailure?: StructuredSubmitFailureDiagnostic; providerMessage?: string; + modelResolution?: ModelResolutionDetail; }; +type ModelResolutionDetail = { kind: string; message: string }; + const PROVIDER_MESSAGE_MAX_CHARS = 300; // The provider's explanation rides on the error context set by the LLM layer; @@ -450,9 +449,12 @@ function providerMessageFromError(error: unknown): string | undefined { return undefined; } const value = error.context?.providerMessage; - if (typeof value !== "string") { - return undefined; - } + return typeof value === "string" ? boundedPublishedText(value) : undefined; +} + +// Scrubbed against the Actions secrets, collapsed to one line, and capped for +// world-readable surfaces (comment, step summary, artifacts, log). +function boundedPublishedText(value: string): string | undefined { const collapsed = scrubGitHubSecrets(value).replace(/\s+/gu, " ").trim(); if (collapsed.length === 0) { return undefined; @@ -462,6 +464,21 @@ function providerMessageFromError(error: unknown): string | undefined { : `${collapsed.slice(0, PROVIDER_MESSAGE_MAX_CHARS - 1).trimEnd()}…`; } +// Model-resolution errors carry codegenie-authored text (with workflow- or +// CLI-supplied model ids) that tells a CI user what to set. Only these error +// messages are published — other error messages can carry external text. +function modelResolutionFromError(error: unknown): ModelResolutionDetail | undefined { + if (!(error instanceof CodegenieError)) { + return undefined; + } + const kind = error.context?.modelResolution; + if (typeof kind !== "string") { + return undefined; + } + const message = boundedPublishedText(error.message); + return message !== undefined ? { kind, message } : undefined; +} + const FAILURE_JSON_MAX_BYTES = 16 * 1024; const FAILURE_MARKDOWN_MAX_BYTES = 4 * 1024; @@ -473,6 +490,7 @@ function publishFailureFiles(input: { errorCode: CodegenieErrorCode | "unknown_error"; diagnostic?: StructuredSubmitFailureDiagnostic; providerMessage?: string; + modelResolution?: ModelResolutionDetail; decision: AuthorizedDecision; runUrl?: string; env: NodeJS.ProcessEnv; @@ -486,15 +504,17 @@ function publishFailureFiles(input: { ...(input.runUrl !== undefined ? { runUrl: input.runUrl } : {}), ...(runId !== undefined && /^\d+$/u.test(runId) ? { runId } : {}), ...(input.diagnostic !== undefined ? { structuredSubmitFailure: input.diagnostic } : {}), - ...(input.providerMessage !== undefined ? { providerMessage: input.providerMessage } : {}) + ...(input.providerMessage !== undefined ? { providerMessage: input.providerMessage } : {}), + ...(input.modelResolution !== undefined ? { modelResolution: input.modelResolution } : {}) }; + const explanation = input.providerMessage ?? input.modelResolution?.message; const json = fitFailureJson(record); const markdown = fitFailureMarkdown([ "# 🧞 Codegenie Review Failed", "", `Error code: \`${input.errorCode}\``, ...(input.diagnostic !== undefined ? ["", renderStructuredSubmitFailure(input.diagnostic)] : []), - ...(input.providerMessage !== undefined ? ["", renderProviderMessage(input.providerMessage)] : []), + ...(explanation !== undefined ? ["", renderProviderMessage(explanation)] : []), ...(input.runUrl !== undefined ? ["", `See the [workflow job](${input.runUrl}) and the failure JSON artifact.`] : []) ].join("\n")); writeFailureFile(input.env.CODEGENIE_FAILURE_PATH, json); @@ -546,26 +566,25 @@ type DecisionRecord = | { eventName: string; run: false; reason: string } | { eventName: string; - run: true; + run: boolean; + reason?: string; lane: AuthorizedDecision["lane"]; prNumber: number; actor: string; association: string; actorAllowlisted: boolean; - permissionCheck: PermissionCheck; - } - | { - eventName: string; - run: false; - reason: string; - lane: AuthorizedDecision["lane"]; - prNumber: number; - actor: string; - association: string; - actorAllowlisted: false; - permissionCheck: "denied"; + permissionCheck: PermissionCheck | "denied"; + modelAlias?: string; + modelSpec?: string; }; +function modelRecordFields(selection: ModelSelection | undefined): { modelAlias?: string; modelSpec?: string } { + return { + ...(selection?.alias !== undefined ? { modelAlias: selection.alias } : {}), + ...(selection !== undefined ? { modelSpec: formatModelSpec(selection.spec) } : {}) + }; +} + function writeDecisionRecord(write: (text: string) => void, record: DecisionRecord): void { write(`github-action: decision ${JSON.stringify(scrubGitHubSecrets(record))}\n`); } @@ -594,6 +613,7 @@ function emitActionRecord( runDir: string | undefined, eventName: string, decision: AuthorizedDecision, + selection: ModelSelection | undefined, outcome: "success" | "review_failed" | "terminal_post_failed", stats: ReturnType["stats"]>, env: NodeJS.ProcessEnv, @@ -609,6 +629,7 @@ function emitActionRecord( association: decision.association, actorAllowlisted: decision.actorAllowlisted, permissionCheck: decision.permissionCheck, + ...modelRecordFields(selection), outcome, ...(errorCode !== undefined ? { errorCode } : {}), runUrl: buildRunUrl(env, env.GITHUB_REPOSITORY ?? "") ?? null, diff --git a/src/github-action/event-gate.ts b/src/github-action/event-gate.ts index 7269b24..63c9ed4 100644 --- a/src/github-action/event-gate.ts +++ b/src/github-action/event-gate.ts @@ -1,7 +1,9 @@ // Pure trigger/authorization decisions over GitHub webhook payloads. No IO: // the live collaborator-permission re-check happens in the entrypoint, this // module only reads the (attacker-visible) payload. Comment text is matched, -// never interpreted — trailing text after the trigger phrase is ignored and +// never interpreted: the only thing read after the trigger phrase is the +// first token of its line, as a candidate alias the entrypoint looks up in +// the workflow's own `models` list (plan 125). Everything else is ignored and // review knobs come exclusively from workflow inputs. export const DEFAULT_TRIGGER_PHRASE = "codegenie review"; @@ -36,6 +38,9 @@ export type TriggerDecision = // Users explicitly allowlisted by workflow input skip the live // write-permission re-check; association-gated actors do not. actorAllowlisted: boolean; + // Comment lane only: lowercased first token after the phrase on its + // line. A lookup key, never a spec — unlisted values are rejected. + requestedAlias?: string; } | { run: false; reason: string }; @@ -52,6 +57,18 @@ export function decideTrigger(eventName: string, payload: unknown, rules: Trigge return skip(`unsupported event: ${eventName}`); } +// The first whitespace-delimited token on the trigger phrase's own line, +// lowercased. Later lines never count, so a comment that continues on the +// next line still means "default". +export function requestedAliasFromComment(body: string, phrase: string): string | undefined { + if (!matchesTriggerPhrase(body, phrase)) { + return undefined; + } + const rest = body.trim().slice(phrase.trim().length); + const token = (rest.split(/\r?\n/u)[0] ?? "").trim().split(/\s+/u)[0] ?? ""; + return token === "" ? undefined : token.toLowerCase(); +} + export function matchesTriggerPhrase(body: string, phrase: string): boolean { const trimmed = body.trim(); const trimmedPhrase = phrase.trim(); @@ -120,7 +137,9 @@ function decideIssueComment(payload: Record, rules: TriggerRule } const actor = stringAt(comment, ["user", "login"]) ?? ""; const association = stringAt(comment, ["author_association"]) ?? "NONE"; - return authorize({ lane: "issue_comment", prNumber, actor, association, isBot: userIsBot(comment) }, rules); + const decision = authorize({ lane: "issue_comment", prNumber, actor, association, isBot: userIsBot(comment) }, rules); + const requestedAlias = requestedAliasFromComment(body, rules.triggerPhrase); + return decision.run && requestedAlias !== undefined ? { ...decision, requestedAlias } : decision; } function authorize( diff --git a/src/github-action/models.ts b/src/github-action/models.ts new file mode 100644 index 0000000..1d4b17d --- /dev/null +++ b/src/github-action/models.ts @@ -0,0 +1,203 @@ +// Model selection for the GitHub Action (plan 125): the `model` / `models` +// inputs, the comment-requested alias, and `llm-api-key` routing. Comment text +// only ever selects an alias the workflow author listed — it never becomes a +// model spec, reasoning level, or flag. +import { isMap, isScalar, LineCounter, parseDocument } from "yaml"; +import { + getCodegeniePiModels, + getPiApiKeyEnvVarName, + getPiCredentialEnvVarNames +} from "../provider/pi-ai-models.js"; +import { splitReasoningSuffix } from "../provider/reasoning.js"; +import { CodegenieError } from "../util/errors.js"; + +export type ModelSpec = { + provider?: string; + model: string; + reasoning: string; +}; + +// The model a run uses; `alias` is set when it was chosen by alias name. +export type ModelSelection = { + alias?: string; + spec: ModelSpec; +}; + +export type ModelConfig = { + // Lowercased alias → spec, in the order the workflow lists them. + aliases: Map; + defaultModel?: ModelSelection; +}; + +const ALIAS_NAME_PATTERN = /^[A-Za-z0-9][A-Za-z0-9._-]{0,31}$/u; + +// One model spec instead of separate provider/model/reasoning inputs: +// `provider/model[:reasoning]`, e.g. `anthropic/claude-opus-5:xhigh`. +// Reasoning defaults to "high" — Action reviews are unattended, so the +// action's posture favors quality over the CLI's interactive default. The +// suffix rule is the shared one (a non-level `:suffix` stays in the model id). +export function parseModelSpec(spec: string): ModelSpec { + const { model: rest, reasoning = "high" } = splitReasoningSuffix(spec); + const slash = rest.indexOf("/"); + const provider = slash > 0 ? rest.slice(0, slash) : undefined; + const model = slash > 0 ? rest.slice(slash + 1) : rest; + if (model.trim() === "" || (slash === 0)) { + throw new CodegenieError("invalid_args", `--model must be provider/model[:reasoning], got: ${spec}`); + } + return { + ...(provider !== undefined ? { provider } : {}), + model, + reasoning + }; +} + +export function formatModelSpec(spec: ModelSpec): string { + return `${spec.provider !== undefined ? `${spec.provider}/` : ""}${spec.model}:${spec.reasoning}`; +} + +// Parses the `models` block: a flat YAML mapping of alias → spec. Errors name +// the line and alias; keys never belong here (see the README), and GitHub +// masks secret values in the log regardless. +export function parseModelAliases( + text: string, + providerExists: (provider: string) => boolean = defaultProviderExists +): Map { + const aliases = new Map(); + if (text.trim() === "") { + return aliases; + } + const lineCounter = new LineCounter(); + // failsafe: every scalar is a string, so aliases like `405:` stay names. + const document = parseDocument(text, { lineCounter, uniqueKeys: false, schema: "failsafe" }); + const [syntaxError] = document.errors; + if (syntaxError !== undefined) { + // Our own message: the yaml library's quotes source text. + throw modelsError(`invalid YAML at line ${lineCounter.linePos(syntaxError.pos[0]).line}`); + } + const contents = document.contents; + if (!isMap(contents)) { + throw modelsError("must be a mapping of alias: provider/model[:reasoning] lines"); + } + for (const pair of contents.items) { + const offset = isScalar(pair.key) ? pair.key.range?.[0] : undefined; + const line = offset !== undefined ? lineCounter.linePos(offset).line : undefined; + const where = line !== undefined ? `line ${line}` : "an entry"; + if (!isScalar(pair.key) || typeof pair.key.value !== "string" || !ALIAS_NAME_PATTERN.test(pair.key.value)) { + throw modelsError(`${where}: alias names must match ${ALIAS_NAME_PATTERN.source}`); + } + const alias = pair.key.value.toLowerCase(); + if (aliases.has(alias)) { + throw modelsError(`${where}: duplicate alias ${alias}`); + } + if (!isScalar(pair.value) || typeof pair.value.value !== "string" || pair.value.value.trim() === "") { + throw modelsError(`${where}: alias ${alias} must map to a provider/model[:reasoning] string`); + } + let spec: ModelSpec; + try { + spec = parseModelSpec(pair.value.value.trim()); + } catch { + throw modelsError(`${where}: alias ${alias} must be provider/model[:reasoning]`); + } + if (spec.provider === undefined) { + throw modelsError(`${where}: alias ${alias} needs a provider prefix (provider/model[:reasoning])`); + } + if (!providerExists(spec.provider)) { + throw modelsError(`${where}: alias ${alias} names unknown provider ${spec.provider}`); + } + aliases.set(alias, spec); + } + return aliases; +} + +// Resolves the `model` input against the aliases. Without `models`, `model` +// parsing is exactly the pre-plan-125 behavior. +export function resolveModelConfig(modelInput: string | undefined, aliases: Map): ModelConfig { + const trimmed = modelInput?.trim() ?? ""; + if (aliases.size === 0) { + return trimmed === "" ? { aliases } : { aliases, defaultModel: { spec: parseModelSpec(trimmed) } }; + } + if (trimmed === "") { + throw new CodegenieError("invalid_args", "--models requires --model to name the default (an alias or provider/model[:reasoning])"); + } + if (!trimmed.includes("/")) { + const alias = trimmed.toLowerCase(); + const spec = aliases.get(alias); + if (spec === undefined) { + throw new CodegenieError("invalid_args", `--model ${trimmed} is not one of the configured model aliases`); + } + return { aliases, defaultModel: { alias, spec } }; + } + return { aliases, defaultModel: { spec: parseModelSpec(trimmed) } }; +} + +// The model for this event. Without aliases, comment text is ignored exactly +// as before; with aliases, an unlisted token is "unknown" (the caller replies). +export function selectModel( + config: ModelConfig, + requestedAlias: string | undefined +): { kind: "selected"; selection?: ModelSelection } | { kind: "unknown_alias" } { + if (config.aliases.size === 0 || requestedAlias === undefined) { + return { kind: "selected", ...(config.defaultModel !== undefined ? { selection: config.defaultModel } : {}) }; + } + const spec = config.aliases.get(requestedAlias.toLowerCase()); + return spec === undefined ? { kind: "unknown_alias" } : { kind: "selected", selection: { alias: requestedAlias.toLowerCase(), spec } }; +} + +// Fixed text built only from configured alias names (restricted charset), so +// nothing from the triggering comment is echoed. +export function renderUnknownAliasReply(config: ModelConfig): string { + const names = [...config.aliases.keys()].map((alias) => ( + alias === config.defaultModel?.alias ? `\`${alias}\` (default)` : `\`${alias}\`` + )); + return `**🧞 Codegenie**: unknown model. Available: ${names.join(", ")}.`; +} + +// `llm-api-key` (LLM_API_KEY) is authoritative when set: every selectable +// model must share one API-key provider, that provider's competing credential +// vars are cleared, and the key is written unconditionally. Unset, provider +// credentials resolve exactly as they always have. Returns the provider the +// key was routed to, if any. +export function applyLlmApiKey(env: NodeJS.ProcessEnv, config: ModelConfig): string | undefined { + const key = env.LLM_API_KEY; + if (key === undefined || key === "") { + return undefined; + } + const selectable = [ + ...(config.defaultModel !== undefined ? [config.defaultModel.spec] : []), + ...config.aliases.values() + ]; + if (selectable.length === 0 || selectable.some((spec) => spec.provider === undefined)) { + throw new CodegenieError( + "invalid_args", + "LLM_API_KEY requires a model input with a provider prefix (e.g. anthropic/claude-opus-5) so the key can be routed" + ); + } + const providers = [...new Set(selectable.map((spec) => spec.provider as string))]; + if (providers.length > 1) { + throw new CodegenieError( + "invalid_args", + `llm-api-key is a single key, but models use providers ${providers.join(", ")}. Remove llm-api-key and set each provider's env var (see models.md#credentials).` + ); + } + const provider = providers[0] as string; + const envVarName = getPiApiKeyEnvVarName(provider); + if (envVarName === undefined) { + throw new CodegenieError( + "invalid_args", + `provider ${provider} does not accept an API key; set its native credentials instead of LLM_API_KEY` + ); + } + for (const name of getPiCredentialEnvVarNames(provider)) { + delete env[name]; + } + env[envVarName] = key; + return provider; +} + +function defaultProviderExists(provider: string): boolean { + return getCodegeniePiModels().getProvider(provider) !== undefined; +} + +function modelsError(detail: string): CodegenieError { + return new CodegenieError("invalid_args", `--models ${detail}`); +} diff --git a/src/github-action/status-comment.ts b/src/github-action/status-comment.ts index 271bb60..1ccf3d1 100644 --- a/src/github-action/status-comment.ts +++ b/src/github-action/status-comment.ts @@ -237,7 +237,10 @@ export function createStatusCommentController(options: StatusCommentOptions): St } terminal = true; await settle(); - const body = appendStatusCommentMarker(reportMarkdown === undefined ? renderFailureBody(errorCode, options.runUrl, diagnostic, providerMessage) + // Both shapes are sanitized: the failure body carries provider prose and + // codegenie error text, which must not ping users or smuggle HTML. + const body = appendStatusCommentMarker(reportMarkdown === undefined + ? sanitizeGitHubCommentBody(renderFailureBody(errorCode, options.runUrl, diagnostic, providerMessage)) : capTerminalBody(sanitizeGitHubCommentBody(reportMarkdown), options.runUrl, STATUS_COMMENT_MARKER.length + 4).body); stats.terminalState = "failure"; const bodyBytes = Buffer.byteLength(body, "utf8"); diff --git a/src/llm/llm-runner.ts b/src/llm/llm-runner.ts index 98e5792..637d9c5 100644 --- a/src/llm/llm-runner.ts +++ b/src/llm/llm-runner.ts @@ -305,8 +305,20 @@ export type PiModelRef = { oauthProvider?: string; }; +// Why resolveModel returned undefined. "unresolved" covers every case the +// resolver cannot attribute (no authenticated provider matched, a provider +// with no usable models, or an unexpected lookup error). +export type ModelResolutionFailure = { + kind: "unknown_model" | "deprecated_model" | "missing_credentials" | "unresolved"; + provider?: string; + model?: string; +}; + export interface PiAiAdapter { resolveModel(input: { provider?: string; model?: string }): PiModelRef | undefined; + // Optional: explains a resolveModel miss. Test adapters may omit it; the + // runner then reports the generic unresolved message. + explainUnresolvedModel?(input: { provider?: string; model?: string }): ModelResolutionFailure; complete( model: PiModelRef, context: { messages: unknown[]; tools: Array<{ name: string; description: string; parameters: TSchema }> }, diff --git a/src/llm/pi-runner.ts b/src/llm/pi-runner.ts index d8d6423..d4ce297 100644 --- a/src/llm/pi-runner.ts +++ b/src/llm/pi-runner.ts @@ -25,7 +25,7 @@ import { import pLimit from "p-limit"; import { createFileAuthStorage, createPiCredentialStore } from "../provider/provider-services.js"; import { filterDeprecatedProviderModels, isDeprecatedProviderModel } from "../provider/model-policy.js"; -import { getCodegeniePiModels, getPiEnvApiKey } from "../provider/pi-ai-models.js"; +import { describeProviderCredentials, getCodegeniePiModels, getPiEnvApiKey } from "../provider/pi-ai-models.js"; import { assertReasoningSupported, modelThinkingLevels, selectReasoningEffort, type ReasoningPolicy } from "../provider/reasoning.js"; import { getCodegeniePaths } from "../config/paths.js"; import { registerSecret, stripCredentials, stripCredentialsWithSummary } from "../telemetry/redaction.js"; @@ -48,6 +48,7 @@ import { type LlmSubmitFailureClassification, type LlmToolResultSummary, type ModelCallCacheMissReason, + type ModelResolutionFailure, type PiAiAdapter, type PiAssistantMessage, type PiModelRef, @@ -217,10 +218,18 @@ export function createPiRunner(opts: CreateRunnerOptions): LlmRunner { model?: string; }); if (!model) { - throw new CodegenieError("config_error", "no usable LLM model could be resolved; run `codegenie provider login ` or configure --provider/--model", { + const request = definedRecord({ provider: opts.llmConfig.provider, model: opts.llmConfig.model }) as { provider?: string; model?: string }; + let failure: ModelResolutionFailure = { kind: "unresolved" }; + try { + failure = adapter.explainUnresolvedModel?.(request) ?? failure; + } catch { + // The explanation is best-effort; the generic message still applies. + } + throw new CodegenieError("config_error", modelResolutionMessage(failure), { context: { provider: opts.llmConfig.provider ?? null, model: opts.llmConfig.model ?? null, + modelResolution: failure.kind, hint: "run `codegenie provider login ` and `codegenie provider models --all` to inspect available authenticated models" } }); @@ -967,6 +976,10 @@ export function createRealPiAiAdapter(deps: RealPiAiAdapterDeps = {}): PiAiAdapt const models = deps.models ?? getCodegeniePiModels(createPiCredentialStore(authStorage)); return { resolveModel: ({ provider, model }) => resolveRealModel(provider, model, authStorage, models), + explainUnresolvedModel: ({ provider, model }) => { + const resolution = resolveRealModelWithReason(provider, model, authStorage, models); + return "failure" in resolution ? resolution.failure : { kind: "unresolved" }; + }, complete: async (model, context, options) => { const { submitToolName, onStreamEvent, onRejectedArguments, ...providerOptions } = options; const streamHooks = { @@ -4169,40 +4182,60 @@ function defaultToolMeta(): ToolResultMeta { return { backend: "text", precision: "text", degraded: false }; } +type ModelResolution = { model: PiModelRef } | { failure: ModelResolutionFailure }; + function resolveRealModel( provider: string | undefined, model: string | undefined, authStorage?: PiAuthStorage, models: Pick = getCodegeniePiModels() ): PiModelRef | undefined { + const resolution = resolveRealModelWithReason(provider, model, authStorage, models); + return "model" in resolution ? resolution.model : undefined; +} + +function resolveRealModelWithReason( + provider: string | undefined, + model: string | undefined, + authStorage?: PiAuthStorage, + models: Pick = getCodegeniePiModels() +): ModelResolution { const qualified = provider === undefined && model ? splitProviderQualifiedModel(model, models) : undefined; const resolvedProvider = provider ?? qualified?.provider; const resolvedModel = qualified?.model ?? model; if (resolvedProvider && resolvedModel) { + const target = { provider: resolvedProvider, model: resolvedModel }; if (isDeprecatedProviderModel(resolvedProvider, resolvedModel)) { - return undefined; + return { failure: { kind: "deprecated_model", ...target } }; } try { const raw = models.getModel(resolvedProvider, resolvedModel); if (!raw) { - return undefined; + return { failure: { kind: "unknown_model", ...target } }; } const auth = resolveProviderAuth(resolvedProvider, authStorage, models); - return auth ? { provider: resolvedProvider, id: resolvedModel, raw: applyModelOverrides(raw), ...auth } : undefined; + return auth + ? { model: { provider: resolvedProvider, id: resolvedModel, raw: applyModelOverrides(raw), ...auth } } + : { failure: { kind: "missing_credentials", ...target } }; } catch { - return undefined; + return { failure: { kind: "unresolved", ...target } }; } } if (resolvedProvider) { + if (models.getProvider(resolvedProvider) === undefined) { + return { failure: { kind: "unresolved", provider: resolvedProvider } }; + } const auth = resolveProviderAuth(resolvedProvider, authStorage, models); if (!auth) { - return undefined; + return { failure: { kind: "missing_credentials", provider: resolvedProvider } }; } const providerModels = filterDeprecatedProviderModels([...models.getModels(resolvedProvider)]); const first = providerModels[0]; - return first ? { provider: resolvedProvider, id: first.id, raw: applyModelOverrides(first), ...auth } : undefined; + return first + ? { model: { provider: resolvedProvider, id: first.id, raw: applyModelOverrides(first), ...auth } } + : { failure: { kind: "unresolved", provider: resolvedProvider } }; } for (const provider of models.getProviders()) { @@ -4214,10 +4247,28 @@ function resolveRealModel( const providerModels = filterDeprecatedProviderModels([...models.getModels(providerId)]); const match = resolvedModel ? providerModels.find((candidate) => candidate.id === resolvedModel) : providerModels[0]; if (match) { - return { provider: providerId, id: match.id, raw: applyModelOverrides(match), ...auth }; + return { model: { provider: providerId, id: match.id, raw: applyModelOverrides(match), ...auth } }; } } - return undefined; + return { failure: { kind: "unresolved", ...(resolvedModel !== undefined ? { model: resolvedModel } : {}) } }; +} + +const UNRESOLVED_MODEL_MESSAGE = "no usable LLM model could be resolved; run `codegenie provider login ` or configure --provider/--model"; + +// Only codegenie-authored text with workflow- or CLI-supplied ids; the +// GitHub Action renders it in the failure comment (scrubbed and sanitized). +export function modelResolutionMessage(failure: ModelResolutionFailure): string { + const target = failure.provider !== undefined && failure.model !== undefined ? `${failure.provider}/${failure.model}` : undefined; + if (failure.kind === "unknown_model" && target !== undefined) { + return `unknown model ${target}; run \`codegenie provider models --all\` to list available models`; + } + if (failure.kind === "deprecated_model" && target !== undefined) { + return `model ${target} is deprecated; choose a current model`; + } + if (failure.kind === "missing_credentials" && failure.provider !== undefined) { + return `no credentials for provider ${failure.provider}: ${describeProviderCredentials(failure.provider)}`; + } + return UNRESOLVED_MODEL_MESSAGE; } function splitProviderQualifiedModel( diff --git a/src/provider/pi-ai-models.ts b/src/provider/pi-ai-models.ts index 62ae4d2..8a79fc2 100644 --- a/src/provider/pi-ai-models.ts +++ b/src/provider/pi-ai-models.ts @@ -6,10 +6,16 @@ import type { CredentialStore, Models, ProviderEnv } from "@earendil-works/pi-ai const piModels = builtinModels(); +// Hand copy of pi-ai's env-api-keys table (not a public export). Keep in sync +// on pi-ai upgrades; the registry coverage test fails when a provider is added +// upstream without a key mapping here or an ambient note below. +// ANTHROPIC_AUTH_TOKEN is deliberately absent: pi's getEnvApiKey skips it too, +// because it is sent as a Bearer header rather than an API key. const API_KEY_ENV_VARS: Record = { "ant-ling": ["ANT_LING_API_KEY"], "anthropic": ["ANTHROPIC_OAUTH_TOKEN", "ANTHROPIC_API_KEY"], "azure-openai-responses": ["AZURE_OPENAI_API_KEY"], + "baseten": ["BASETEN_API_KEY"], "cerebras": ["CEREBRAS_API_KEY"], "cloudflare-ai-gateway": ["CLOUDFLARE_API_KEY"], "cloudflare-workers-ai": ["CLOUDFLARE_API_KEY"], @@ -21,6 +27,7 @@ const API_KEY_ENV_VARS: Record = { "groq": ["GROQ_API_KEY"], "huggingface": ["HF_TOKEN"], "kimi-coding": ["KIMI_API_KEY"], + "meta": ["META_API_KEY"], "minimax": ["MINIMAX_API_KEY"], "minimax-cn": ["MINIMAX_CN_API_KEY"], "mistral": ["MISTRAL_API_KEY"], @@ -31,6 +38,10 @@ const API_KEY_ENV_VARS: Record = { "opencode-go": ["OPENCODE_API_KEY"], "openai": ["OPENAI_API_KEY"], "openrouter": ["OPENROUTER_API_KEY"], + "qwen-token-plan": ["QWEN_TOKEN_PLAN_API_KEY"], + "qwen-token-plan-cn": ["QWEN_TOKEN_PLAN_CN_API_KEY"], + "qwen-token-plan-individual": ["QWEN_TOKEN_PLAN_API_KEY"], + "radius": ["RADIUS_API_KEY"], "together": ["TOGETHER_API_KEY"], "vercel-ai-gateway": ["AI_GATEWAY_API_KEY"], "xai": ["XAI_API_KEY"], @@ -42,6 +53,19 @@ const API_KEY_ENV_VARS: Record = { "zai-coding-cn": ["ZAI_CODING_CN_API_KEY"] }; +// Bearer-token vars pi reads outside the API-key lookup above. Only used to +// clear competing credentials when the Action's llm-api-key is authoritative. +const BEARER_TOKEN_ENV_VARS: Record = { + "anthropic": ["ANTHROPIC_AUTH_TOKEN"] +}; + +// Providers whose credentials are not (only) a single API-key env var. +const AMBIENT_CREDENTIAL_NOTES: Record = { + "amazon-bedrock": "AWS credentials: `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY`, `AWS_BEARER_TOKEN_BEDROCK`, `AWS_PROFILE`, or an OIDC web identity (e.g. aws-actions/configure-aws-credentials)", + "google-vertex": "or Application Default Credentials plus `GOOGLE_CLOUD_PROJECT` and `GOOGLE_CLOUD_LOCATION` (e.g. google-github-actions/auth)", + "openai-codex": "stored ChatGPT-plan login only (`codegenie provider login openai-codex`); in CI this needs a self-hosted runner with that login" +}; + export function getCodegeniePiModels(credentials?: CredentialStore): Models { return credentials === undefined ? piModels : builtinModels({ credentials }); } @@ -57,6 +81,31 @@ export function getPiApiKeyEnvVarName(provider: string): string | undefined { return envVars.find((name) => name.endsWith("_API_KEY")) ?? envVars[envVars.length - 1]; } +// Every API-key env var pi reads for a provider, in lookup order. +export function getPiApiKeyEnvVarNames(provider: string): string[] { + return [...(API_KEY_ENV_VARS[provider] ?? [])]; +} + +// Every env var that can authenticate a provider: the API-key lookup names +// plus bearer-token vars. Clearing these leaves only an explicitly routed key. +export function getPiCredentialEnvVarNames(provider: string): string[] { + return [...(BEARER_TOKEN_ENV_VARS[provider] ?? []), ...(API_KEY_ENV_VARS[provider] ?? [])]; +} + +export function getAmbientCredentialNote(provider: string): string | undefined { + return AMBIENT_CREDENTIAL_NOTES[provider]; +} + +// One-line "how to authenticate" hint for error messages. +export function describeProviderCredentials(provider: string): string { + const envVar = getPiApiKeyEnvVarName(provider); + const note = getAmbientCredentialNote(provider); + if (envVar !== undefined) { + return `set ${envVar}${note !== undefined ? ` ${note}` : ""} (or run \`codegenie provider login ${provider}\`)`; + } + return note ?? `run \`codegenie provider login ${provider}\``; +} + export function getPiEnvApiKey(provider: string, env?: ProviderEnv): string | undefined { const envVars = API_KEY_ENV_VARS[provider]; const firstKey = envVars?.find((name) => providerEnvValue(name, env) !== undefined); diff --git a/tests/github-action.test.ts b/tests/github-action.test.ts index 0b25559..2423348 100644 --- a/tests/github-action.test.ts +++ b/tests/github-action.test.ts @@ -8,10 +8,11 @@ import { DEFAULT_ALLOWED_ASSOCIATIONS, DEFAULT_TRIGGER_PHRASE, decideTrigger, - matchesTriggerPhrase + matchesTriggerPhrase, + requestedAliasFromComment } from "../src/github-action/event-gate.js"; import { - applyGenericApiKey, + applyLlmApiKey, executeGitHubActionCommand, parseGitHubActionArgs, parseModelSpec, @@ -21,6 +22,12 @@ import type { IssueComment, IssueCommentClient } from "../src/github-action/issu import { createIssueCommentClient } from "../src/github-action/issue-comments.js"; import { appendStatusCommentMarker, STATUS_COMMENT_MARKER } from "../src/github-action/marker.js"; import { createStatusCommentController } from "../src/github-action/status-comment.js"; +import { + parseModelAliases, + renderUnknownAliasReply, + resolveModelConfig, + selectModel +} from "../src/github-action/models.js"; import { ISSUE_COMMENT_MAX_CHARS, TRUNCATION_DISCLOSURE } from "../src/github-action/render.js"; import { createGitHubClient } from "../src/github/github-client.js"; import type { runGh } from "../src/git/subprocess.js"; @@ -156,6 +163,125 @@ describe("github-action event gate", () => { expect(matchesTriggerPhrase("I think codegenie review is neat", "codegenie review")).toBe(false); expect(matchesTriggerPhrase("anything", "")).toBe(false); }); + + it("extracts a requested alias from the phrase's own line only", () => { + expect(requestedAliasFromComment("codegenie review opus", "codegenie review")).toBe("opus"); + expect(requestedAliasFromComment(" codegenie review OPUS please\n", "codegenie review")).toBe("opus"); + expect(requestedAliasFromComment("codegenie review\tglm", "codegenie review")).toBe("glm"); + expect(requestedAliasFromComment("codegenie review", "codegenie review")).toBeUndefined(); + expect(requestedAliasFromComment("codegenie review\n\nfocus on auth", "codegenie review")).toBeUndefined(); + expect(requestedAliasFromComment("codegenie review \r\nopus", "codegenie review")).toBeUndefined(); + expect(requestedAliasFromComment("codegenie reviewopus", "codegenie review")).toBeUndefined(); + expect(requestedAliasFromComment("please codegenie review opus", "codegenie review")).toBeUndefined(); + + expect(decideTrigger("issue_comment", issueCommentPayload({ body: "codegenie review GLM" }), RULES)).toMatchObject({ + run: true, + requestedAlias: "glm" + }); + expect(decideTrigger("issue_comment", issueCommentPayload(), RULES)).not.toHaveProperty("requestedAlias"); + expect(decideTrigger("pull_request", pullRequestPayload(), RULES)).not.toHaveProperty("requestedAlias"); + }); +}); + +const MODELS = [ + "luna: openrouter/openai/gpt-6-luna:xhigh", + "deepseek: openrouter/deepseek/deepseek-v4.1-flash:max", + "opus: anthropic/claude-opus-5" +].join("\n"); + +describe("github-action model aliases", () => { + it("parses a flat alias block with comments, quotes, blank lines and mixed case", () => { + const aliases = parseModelAliases([ + "# models for this repo", + "Luna: openrouter/openai/gpt-6-luna:xhigh", + "", + "\"opus\": \"anthropic/claude-opus-5\" # default reasoning", + "free: openrouter/cohere/north-mini-code:free", + "freehigh: openrouter/cohere/north-mini-code:free:high", + "freelow: openrouter/cohere/north-mini-code:free:low", + "typo: anthropic/claude-opus-5:hgh" + ].join("\n")); + expect([...aliases.keys()]).toEqual(["luna", "opus", "free", "freehigh", "freelow", "typo"]); + expect(aliases.get("luna")).toEqual({ provider: "openrouter", model: "openai/gpt-6-luna", reasoning: "xhigh" }); + expect(aliases.get("opus")).toEqual({ provider: "anthropic", model: "claude-opus-5", reasoning: "high" }); + // Only recognized reasoning suffixes split; other suffixes stay in the id. + expect(aliases.get("free")).toEqual({ provider: "openrouter", model: "cohere/north-mini-code:free", reasoning: "high" }); + expect(aliases.get("freehigh")).toEqual({ provider: "openrouter", model: "cohere/north-mini-code:free", reasoning: "high" }); + expect(aliases.get("freelow")).toEqual({ provider: "openrouter", model: "cohere/north-mini-code:free", reasoning: "low" }); + // A typo'd level is left for selected-model lookup, not rejected here. + expect(aliases.get("typo")).toEqual({ provider: "anthropic", model: "claude-opus-5:hgh", reasoning: "high" }); + expect(parseModelAliases(" \n").size).toBe(0); + // Numeric-looking names are names, not YAML numbers. + expect([...parseModelAliases("405: openrouter/meta-llama/llama-3.1-405b-instruct\n1.5: anthropic/claude-opus-5").keys()]) + .toEqual(["405", "1.5"]); + }); + + it("rejects malformed alias blocks with line-level errors that never echo values", () => { + const cases: Array<[string, RegExp]> = [ + ["Opus: anthropic/claude-opus-5\nopus: anthropic/claude-sonnet-5", /line 2: duplicate alias opus/u], + ["-bad: anthropic/claude-opus-5", /line 1: alias names must match/u], + ["has space: anthropic/claude-opus-5", /line 1: alias names must match/u], + ["opus:\n model: anthropic/claude-opus-5", /line 1: alias opus must map to a provider\/model/u], + ["opus: [anthropic/claude-opus-5]", /line 1: alias opus must map to a provider\/model/u], + ["base: &spec anthropic/claude-opus-5\nopus: *spec", /line 2: alias opus must map to a provider\/model/u], + ["- anthropic/claude-opus-5", /must be a mapping/u], + ["opus: claude-opus-5", /line 1: alias opus needs a provider prefix/u], + ["opus: nope/claude-opus-5", /line 1: alias opus names unknown provider nope/u], + ["opus: /claude-opus-5", /line 1: alias opus must be provider\/model/u] + ]; + for (const [text, pattern] of cases) { + expect(() => parseModelAliases(text), text).toThrow(pattern); + } + let caught: unknown; + try { + parseModelAliases("ok: anthropic/claude-opus-5\nbad: \"sk-ant-unterminated-secret"); + } catch (error) { + caught = error; + } + expect(caught).toMatchObject({ code: "invalid_args", message: expect.stringMatching(/invalid YAML at line 2/u) }); + expect(String((caught as Error).message)).not.toContain("sk-ant-unterminated-secret"); + }); + + it("resolves the default from an alias or a full spec, and keeps the no-models behavior", () => { + const aliases = parseModelAliases(MODELS); + expect(resolveModelConfig("LUNA", aliases).defaultModel).toEqual({ + alias: "luna", + spec: { provider: "openrouter", model: "openai/gpt-6-luna", reasoning: "xhigh" } + }); + expect(resolveModelConfig("openai/gpt-5.5", aliases).defaultModel).toEqual({ + spec: { provider: "openai", model: "gpt-5.5", reasoning: "high" } + }); + expect(() => resolveModelConfig("sonnet", aliases)).toThrow(/not one of the configured model aliases/u); + expect(() => resolveModelConfig("", aliases)).toThrow(/--models requires --model/u); + // Without models, `model` parses exactly as before (provider-less allowed). + expect(resolveModelConfig("opus", new Map()).defaultModel).toEqual({ spec: { model: "opus", reasoning: "high" } }); + expect(resolveModelConfig(" ", new Map()).defaultModel).toBeUndefined(); + }); + + it("selects the requested alias, ignores tokens without models, and lists names for unknowns", () => { + const config = resolveModelConfig("luna", parseModelAliases(MODELS)); + expect(selectModel(config, undefined)).toEqual({ kind: "selected", selection: config.defaultModel }); + expect(selectModel(config, "OPUS")).toMatchObject({ kind: "selected", selection: { alias: "opus" } }); + expect(selectModel(config, "opsu")).toEqual({ kind: "unknown_alias" }); + const legacy = resolveModelConfig("openrouter/deepseek/deepseek-v4.1-flash:max", new Map()); + expect(selectModel(legacy, "opsu")).toEqual({ kind: "selected", selection: legacy.defaultModel }); + expect(renderUnknownAliasReply(config)).toBe( + "**🧞 Codegenie**: unknown model. Available: `luna` (default), `deepseek`, `opus`." + ); + expect(renderUnknownAliasReply(resolveModelConfig("openai/gpt-5.5", parseModelAliases(MODELS)))).toBe( + "**🧞 Codegenie**: unknown model. Available: `luna`, `deepseek`, `opus`." + ); + }); + + it("validates model and models together when parsing action flags", () => { + expect(() => parseGitHubActionArgs(["--models", MODELS])).toThrow(/--models requires --model/u); + expect(() => parseGitHubActionArgs(["--models", MODELS, "--model", "sonnet"])).toThrow(/not one of the configured/u); + const parsed = parseGitHubActionArgs(["--models", MODELS, "--model", "luna"]); + expect(parsed.models.defaultModel?.alias).toBe("luna"); + expect(parsed.models.aliases.size).toBe(3); + // The composite action always forwards both flags, possibly empty. + expect(parseGitHubActionArgs(["--model", "", "--models", ""]).models).toEqual({ aliases: new Map() }); + }); }); type FakeCommentCall = @@ -1088,22 +1214,63 @@ describe("github-action entrypoint", () => { expect(result.reportMarkdown).not.toContain("concise posting summary"); }); - it("routes LLM_API_KEY to the provider's env var without clobbering native vars", () => { + it("makes LLM_API_KEY the only key for its provider, overriding native vars", () => { + const single = (model: string) => resolveModelConfig(model, new Map()); const env: NodeJS.ProcessEnv = { LLM_API_KEY: "generic-key" }; - applyGenericApiKey(env, parseModelSpec("anthropic/claude-opus-4-8")); + applyLlmApiKey(env, single("anthropic/claude-opus-4-8")); expect(env.ANTHROPIC_API_KEY).toBe("generic-key"); + // Plan 125 precedence: an explicit llm-api-key beats a pre-set native var. const preset: NodeJS.ProcessEnv = { LLM_API_KEY: "generic-key", OPENAI_API_KEY: "native-key" }; - applyGenericApiKey(preset, parseModelSpec("openai/gpt-5.5")); - expect(preset.OPENAI_API_KEY).toBe("native-key"); + applyLlmApiKey(preset, single("openai/gpt-5.5")); + expect(preset.OPENAI_API_KEY).toBe("generic-key"); + + // Competing credentials for the same provider are cleared; others stay. + const competing: NodeJS.ProcessEnv = { + LLM_API_KEY: "generic-key", + ANTHROPIC_AUTH_TOKEN: "bearer", + ANTHROPIC_OAUTH_TOKEN: "oauth", + OPENAI_API_KEY: "unrelated" + }; + applyLlmApiKey(competing, single("anthropic/claude-opus-5")); + expect(competing).toEqual({ LLM_API_KEY: "generic-key", ANTHROPIC_API_KEY: "generic-key", OPENAI_API_KEY: "unrelated" }); - expect(() => applyGenericApiKey({ LLM_API_KEY: "k" }, parseModelSpec("opus"))).toThrow(/provider prefix/u); - expect(() => applyGenericApiKey({ LLM_API_KEY: "k" }, parseModelSpec("not-a-provider/x"))).toThrow( + expect(() => applyLlmApiKey({ LLM_API_KEY: "k" }, single("opus"))).toThrow(/provider prefix/u); + expect(() => applyLlmApiKey({ LLM_API_KEY: "k" }, { aliases: new Map() })).toThrow(/provider prefix/u); + expect(() => applyLlmApiKey({ LLM_API_KEY: "k" }, single("not-a-provider/x"))).toThrow(/does not accept an API key/u); + expect(() => applyLlmApiKey({ LLM_API_KEY: "k" }, single("amazon-bedrock/anthropic.claude-sonnet-5"))).toThrow( /does not accept an API key/u ); - const untouched: NodeJS.ProcessEnv = {}; - applyGenericApiKey(untouched, undefined); - expect(untouched).toEqual({}); + const untouched: NodeJS.ProcessEnv = { OPENAI_API_KEY: "native-key" }; + applyLlmApiKey(untouched, single("openai/gpt-5.5")); + expect(untouched).toEqual({ OPENAI_API_KEY: "native-key" }); + }); + + it("requires every configured model to share llm-api-key's provider", () => { + const secret = "sk-or-v1-MUSTNOTSURFACE"; + const mixed = resolveModelConfig("luna", parseModelAliases([ + "luna: openrouter/openai/gpt-6-luna:xhigh", + "opus: anthropic/claude-opus-5" + ].join("\n"))); + let caught: unknown; + try { + applyLlmApiKey({ LLM_API_KEY: secret }, mixed); + } catch (error) { + caught = error; + } + expect(caught).toMatchObject({ + code: "invalid_args", + message: expect.stringContaining("llm-api-key is a single key, but models use providers openrouter, anthropic") + }); + expect(String((caught as Error).message)).not.toContain(secret); + + const sameProvider = resolveModelConfig("luna", parseModelAliases([ + "luna: openrouter/openai/gpt-6-luna:xhigh", + "glm: openrouter/z-ai/glm-5.3:max" + ].join("\n"))); + const env: NodeJS.ProcessEnv = { LLM_API_KEY: secret }; + applyLlmApiKey(env, sameProvider); + expect(env.OPENROUTER_API_KEY).toBe(secret); }); it("expands the model spec into the synthesized review argv", async () => { @@ -1125,6 +1292,221 @@ describe("github-action entrypoint", () => { ]); }); + async function runWithModels( + payload: Record, + eventName: string, + args: string[], + extra: { env?: Record; comments?: ReturnType } = {} + ): Promise<{ argv?: string[]; output: string; calls: FakeCommentCall[]; env: NodeJS.ProcessEnv }> { + const fake = extra.comments ?? createFakeComments(); + const env = actionEnv(payload, eventName, extra.env); + let argv: string[] | undefined; + let output = ""; + await executeGitHubActionCommand(args, { + env, + issueComments: fake.client, + authStorage: { get: () => undefined }, + minEditIntervalMs: 0, + writeOutput: (text) => { + output += text; + }, + runReview: async (reviewArgv) => { + argv = reviewArgv; + return { runId: "r1", runDir: "", reportMarkdown: "# report" }; + } + }); + return { ...(argv !== undefined ? { argv } : {}), output, calls: fake.calls, env }; + } + + const MODEL_ARGS = ["--model", "luna", "--models", MODELS]; + + it("reviews with the default on the PR lane and a bare comment, and with the alias when requested", async () => { + const pr = await runWithModels(pullRequestPayload(), "pull_request", MODEL_ARGS); + expect(pr.argv?.slice(-6)).toEqual(["--provider", "openrouter", "--model", "openai/gpt-6-luna", "--reasoning", "xhigh"]); + expect(pr.output).toContain('"modelAlias":"luna"'); + expect(pr.output).toContain('"modelSpec":"openrouter/openai/gpt-6-luna:xhigh"'); + + const bare = await runWithModels(issueCommentPayload(), "issue_comment", MODEL_ARGS); + expect(bare.argv?.slice(-6)).toEqual(["--provider", "openrouter", "--model", "openai/gpt-6-luna", "--reasoning", "xhigh"]); + + const opus = await runWithModels(issueCommentPayload({ body: "codegenie review OPUS\nplease focus on auth" }), "issue_comment", MODEL_ARGS); + expect(opus.argv?.slice(-6)).toEqual(["--provider", "anthropic", "--model", "claude-opus-5", "--reasoning", "high"]); + expect(opus.output).toContain('"modelAlias":"opus"'); + }); + + it("replies once to an unknown alias without claiming the status comment or reviewing", async () => { + const result = await runWithModels(issueCommentPayload({ body: "codegenie review opsu @someone" }), "issue_comment", MODEL_ARGS); + expect(result.argv).toBeUndefined(); + expect(result.calls).toEqual([ + { kind: "permission", login: "alice" }, + { kind: "create", issueNumber: 7, body: "**🧞 Codegenie**: unknown model. Available: `luna` (default), `deepseek`, `opus`." } + ]); + expect(result.output).toContain("skipped — unknown model alias"); + expect(result.output).toContain('"reason":"unknown model alias"'); + }); + + it("stays silent for unauthorized commenters, even with an unknown alias", async () => { + const denied = createFakeComments({ permission: "read" }); + const result = await runWithModels(issueCommentPayload({ body: "codegenie review opsu" }), "issue_comment", MODEL_ARGS, { comments: denied }); + expect(result.argv).toBeUndefined(); + expect(result.calls).toEqual([{ kind: "permission", login: "alice" }]); + }); + + it("resolves aliases in preflight without model credentials", async () => { + async function preflight(body: string): Promise<{ outputs: Record; calls: FakeCommentCall[] }> { + const fake = createFakeComments(); + const outputPath = path.join(scratch, `alias-output-${Math.random().toString(36).slice(2)}.txt`); + await executeGitHubActionCommand(["--preflight-only", "true", ...MODEL_ARGS], { + env: actionEnv(issueCommentPayload({ body }), "issue_comment", { GITHUB_OUTPUT: outputPath, GITHUB_ACTIONS: "true" }), + issueComments: fake.client, + writeOutput: () => undefined, + runReview: async () => { + throw new Error("preflight must not run a review"); + } + }); + const outputs = Object.fromEntries( + readFileSync(outputPath, "utf8").trim().split("\n").map((line) => line.split("=", 2) as [string, string]) + ); + return { outputs, calls: fake.calls }; + } + const unknown = await preflight("codegenie review opsu"); + expect(unknown.outputs).toEqual({ "should-run": "false" }); + expect(unknown.calls.filter((call) => call.kind === "create")).toHaveLength(1); + + const known = await preflight("codegenie review opus"); + expect(known.outputs).toEqual({ "should-run": "true", "pr-number": "7" }); + expect(known.calls).toEqual([{ kind: "permission", login: "alice" }]); + }); + + it("uses each alias's own provider env var when llm-api-key is unset", async () => { + const extra = { env: { OPENROUTER_API_KEY: "or-native", ANTHROPIC_API_KEY: "ant-native" } }; + const opus = await runWithModels(issueCommentPayload({ body: "codegenie review opus" }), "issue_comment", MODEL_ARGS, extra); + expect(opus.argv).toContain("claude-opus-5"); + expect(opus.env.ANTHROPIC_API_KEY).toBe("ant-native"); + expect(opus.env.OPENROUTER_API_KEY).toBe("or-native"); + }); + + it("fails a multi-provider alias set given one llm-api-key before claiming the status comment", async () => { + const fake = createFakeComments(); + await expect( + executeGitHubActionCommand(MODEL_ARGS, { + env: actionEnv(pullRequestPayload(), "pull_request", { LLM_API_KEY: "sk-or-v1-one-key" }), + issueComments: fake.client, + writeOutput: () => undefined, + runReview: async () => { + throw new Error("review must not run"); + } + }) + ).rejects.toThrow(/llm-api-key is a single key, but models use providers openrouter, anthropic/u); + expect(fake.calls.some((call) => call.kind === "create" || call.kind === "update")).toBe(false); + }); + + it("refuses llm-api-key when a stored login on the runner would override it", async () => { + const fake = createFakeComments(); + const stored = { type: "api_key" as const, apiKey: "sk-or-stored", createdAt: new Date(0).toISOString() }; + await expect( + executeGitHubActionCommand(["--model", "openrouter/deepseek/deepseek-v4.1-flash:max"], { + env: actionEnv(pullRequestPayload(), "pull_request", { LLM_API_KEY: "sk-or-v1-explicit" }), + issueComments: fake.client, + authStorage: { get: (provider) => (provider === "openrouter" ? stored : undefined) }, + writeOutput: () => undefined, + runReview: async () => { + throw new Error("review must not run"); + } + }) + ).rejects.toThrow(/stored codegenie login for openrouter on this runner would override llm-api-key/u); + expect(fake.calls.some((call) => call.kind === "create" || call.kind === "update")).toBe(false); + + // A stored login for another provider, or no llm-api-key, is unaffected. + const other = await runWithModels(pullRequestPayload(), "pull_request", ["--model", "anthropic/claude-opus-5"], { + env: { LLM_API_KEY: "sk-ant-explicit" } + }); + expect(other.env.ANTHROPIC_API_KEY).toBe("sk-ant-explicit"); + }); + + // The README's simple form must behave exactly as it did before plan 125. + it("keeps the simple single-model form backward compatible", async () => { + const simple = ["--model", "openrouter/deepseek/deepseek-v4.1-flash:max", "--models", ""]; + const expectedArgv = [ + "review", "--pr", "7", "--ci", "--post-github-comments", + "--provider", "openrouter", "--model", "deepseek/deepseek-v4.1-flash", "--reasoning", "max" + ]; + const withKey = await runWithModels( + issueCommentPayload({ body: "codegenie review opus please" }), + "issue_comment", + simple, + { env: { LLM_API_KEY: "sk-or-v1-simple" } } + ); + expect(withKey.argv).toEqual(expectedArgv); + expect(withKey.env.OPENROUTER_API_KEY).toBe("sk-or-v1-simple"); + expect(withKey.calls.filter((call) => call.kind === "create")).toHaveLength(1); // the status comment only + + const nativeOnly = await runWithModels(pullRequestPayload(), "pull_request", simple, { env: { OPENROUTER_API_KEY: "or-native" } }); + expect(nativeOnly.argv).toEqual(expectedArgv); + expect(nativeOnly.env.OPENROUTER_API_KEY).toBe("or-native"); + }); + + it("shows model-resolution messages in the failure comment, and only those", async () => { + async function fail(error: CodegenieError): Promise<{ body: string; failure: Record }> { + const fake = createFakeComments(); + const failurePath = path.join(scratch, `resolution-${Math.random().toString(36).slice(2)}.json`); + await expect( + executeGitHubActionCommand(MODEL_ARGS, { + env: actionEnv(pullRequestPayload(), "pull_request", { CODEGENIE_FAILURE_PATH: failurePath }), + issueComments: fake.client, + minEditIntervalMs: 0, + writeOutput: () => undefined, + runReview: async () => { + throw error; + } + }) + ).rejects.toBe(error); + const terminal = fake.calls.at(-1) as { kind: string; body: string }; + return { body: terminal.body, failure: JSON.parse(readFileSync(failurePath, "utf8")) as Record }; + } + + const message = "no credentials for provider anthropic: set ANTHROPIC_API_KEY (or run `codegenie provider login anthropic`)"; + const missing = await fail(new CodegenieError("config_error", message, { context: { modelResolution: "missing_credentials" } })); + expect(missing.body).toContain("`config_error`"); + expect(missing.body).toContain(message); + expect(missing.failure).toMatchObject({ errorCode: "config_error", modelResolution: { kind: "missing_credentials", message } }); + expect(missing.failure).not.toHaveProperty("providerMessage"); + + const generic = await fail(new CodegenieError("config_error", "git stderr: fatal: /home/runner/private/path")); + expect(generic.body).toContain("`config_error`"); + expect(generic.body).not.toContain("/home/runner/private/path"); + expect(generic.failure).not.toHaveProperty("modelResolution"); + }); + + it("sanitizes every failure comment: mentions, HTML comments and secrets", async () => { + async function failureBody(error: CodegenieError): Promise { + const fake = createFakeComments(); + await expect( + executeGitHubActionCommand([], { + env: actionEnv(pullRequestPayload(), "pull_request"), + issueComments: fake.client, + minEditIntervalMs: 0, + writeOutput: () => undefined, + runReview: async () => { + throw error; + } + }) + ).rejects.toBe(error); + return (fake.calls.at(-1) as { body: string }).body; + } + const hostile = "ping @octocat token=abcdefghijklmnopqrstuv"; + for (const error of [ + new CodegenieError("llm_call_failed", "provider failed", { context: { providerMessage: hostile } }), + new CodegenieError("config_error", `unknown model openrouter/x; ${hostile}`, { context: { modelResolution: "unknown_model" } }) + ]) { + const body = await failureBody(error); + expect(body).toContain("`@octocat`"); + expect(body).not.toContain(""); + expect(body).not.toContain("abcdefghijklmnopqrstuv"); + expect(body).toContain(STATUS_COMMENT_MARKER); + } + }); + it("rejects unknown flags and invalid booleans", () => { expect(() => parseGitHubActionArgs(["--bogus", "x"])).toThrow(/unknown github-action flag/u); expect(() => parseGitHubActionArgs(["--on-pull-request", "yes"])).toThrow(/must be/u); @@ -1227,6 +1609,7 @@ type WorkflowJob = { }; type WorkflowDocument = { + on?: Record; concurrency?: { group?: string; "cancel-in-progress"?: boolean | string }; jobs: Record; }; @@ -1234,8 +1617,7 @@ type WorkflowDocument = { describe("GitHub Action and workflow contracts", () => { const workflowPaths = [ ".github/workflows/codegenie-review.yml", - "examples/workflows/codegenie-review-comment.yml", - "examples/workflows/codegenie-review-pr.yml" + "examples/workflows/codegenie-review.yml" ]; it("forwards preflight and bot identity inputs through the composite action", () => { @@ -1254,6 +1636,9 @@ describe("GitHub Action and workflow contracts", () => { const runStep = action.runs.steps.find((step) => step.id === "run"); expect(runStep?.run).toContain('args+=(--bot-login "$INPUT_BOT_LOGIN")'); + expect(action.inputs.models).toBeDefined(); + expect(runStep?.env?.INPUT_MODELS).toBe("${{ inputs.models }}"); + expect(runStep?.run).toContain('args+=(--models "$INPUT_MODELS")'); expect(runStep?.run).toContain('args+=(--preflight-only "$INPUT_PREFLIGHT_ONLY")'); const failurePath = runStep?.env?.CODEGENIE_FAILURE_PATH; expect(failurePath).toBe("${{ runner.temp }}/codegenie-failure.json"); @@ -1278,9 +1663,18 @@ describe("GitHub Action and workflow contracts", () => { expect(workflow.concurrency?.["cancel-in-progress"], workflowPath).toBe(true); return workflow; }); - const commentExample = documents[1]; - const commentReviewStep = commentExample?.jobs.review?.steps?.find((step) => step.uses?.startsWith("0xPolygon/codegenie@")); - expect(commentReviewStep?.with?.["trigger-phrase"]).toBe("codegenie review"); + const example = documents[1]; + expect(Object.keys(example?.on ?? {}).sort()).toEqual(["issue_comment", "pull_request"]); + const exampleStep = example?.jobs.review?.steps?.find((step) => step.uses?.startsWith("0xPolygon/codegenie@")); + expect(exampleStep?.with?.["trigger-phrase"]).toBe("codegenie review"); + // The example's alias list must parse with the real parser, and its + // default must be one of the aliases. + const config = resolveModelConfig(exampleStep?.with?.model, parseModelAliases(exampleStep?.with?.models ?? "")); + expect(config.aliases.size).toBeGreaterThan(1); + expect(config.defaultModel?.alias).toBe(exampleStep?.with?.model); + // models holds specs only; keys go in the step env. + expect(exampleStep?.with?.models).not.toContain("secrets."); + expect(Object.keys(exampleStep?.env ?? {})).toEqual(["OPENROUTER_API_KEY", "ANTHROPIC_API_KEY"]); }); // The action installs the npm version read from its own package.json, so a @@ -1302,16 +1696,14 @@ describe("GitHub Action and workflow contracts", () => { } }); - it("pins pull-request jobs to the base SHA and leaves comment jobs on the default branch", () => { - const dogfood = parseYaml(readFileSync(path.resolve(workflowPaths[0] ?? ""), "utf8")) as WorkflowDocument; - const prExample = parseYaml(readFileSync(path.resolve(workflowPaths[2] ?? ""), "utf8")) as WorkflowDocument; - const commentExample = parseYaml(readFileSync(path.resolve(workflowPaths[1] ?? ""), "utf8")) as WorkflowDocument; - const dogfoodCheckout = dogfood.jobs.review?.steps?.find((step) => step.uses === "actions/checkout@v7"); - expect(dogfoodCheckout?.with?.ref).toContain("github.event.pull_request.base.sha || ''"); - const prCheckout = prExample.jobs.review?.steps?.find((step) => step.uses === "actions/checkout@v7"); - expect(prCheckout?.with?.ref).toContain("github.event.pull_request.base.sha"); - const commentCheckout = commentExample.jobs.review?.steps?.find((step) => step.uses === "actions/checkout@v7"); - expect(commentCheckout?.with?.ref).toBeUndefined(); + it("pins pull-request runs to the base SHA and leaves comment runs on the default branch", () => { + for (const workflowPath of workflowPaths) { + const workflow = parseYaml(readFileSync(path.resolve(workflowPath), "utf8")) as WorkflowDocument; + const checkout = workflow.jobs.review?.steps?.find((step) => step.uses === "actions/checkout@v7"); + // Evaluates to "" on issue_comment, so checkout falls back to the default branch. + expect(checkout?.with?.ref, workflowPath).toContain("github.event.pull_request.base.sha || ''"); + expect(workflow.concurrency?.group, workflowPath).toContain("github.event.pull_request.number || github.event.issue.number"); + } }); it("runs standalone CI against the PR head with actionlint available before every gate", () => { diff --git a/tests/model-resolution.test.ts b/tests/model-resolution.test.ts new file mode 100644 index 0000000..2bba1bf --- /dev/null +++ b/tests/model-resolution.test.ts @@ -0,0 +1,187 @@ +import type { Models } from "@earendil-works/pi-ai"; +import { afterEach, describe, expect, it, vi } from "vitest"; +import { applyLlmApiKey, parseModelAliases, resolveModelConfig } from "../src/github-action/models.js"; +import type { PiAiAdapter } from "../src/llm/llm-runner.js"; +import { createPiRunner, createRealPiAiAdapter, modelResolutionMessage } from "../src/llm/pi-runner.js"; +import { createPiModelRegistry, type PiAuthStorage, type ProviderAuthEntry } from "../src/provider/provider-services.js"; +import { + describeProviderCredentials, + getAmbientCredentialNote, + getPiApiKeyEnvVarNames, + getPiCredentialEnvVarNames +} from "../src/provider/pi-ai-models.js"; +import type { TelemetryRecorder } from "../src/telemetry/telemetry-recorder.js"; +import type { Logger } from "../src/types.js"; +import { CodegenieError } from "../src/util/errors.js"; + +function authStorage(entries: Record = {}): PiAuthStorage { + return { + loadAll: () => ({ ...entries }), + get: (provider) => entries[provider], + set: () => undefined, + delete: () => undefined, + clear: () => undefined + }; +} + +// providerEnvValue falls back to process.env, so every credential var a test +// relies on being absent is stubbed empty (falsy) rather than assumed unset. +function clearProviderEnv(provider: string): void { + for (const name of getPiCredentialEnvVarNames(provider)) { + vi.stubEnv(name, ""); + } +} + +afterEach(() => { + vi.unstubAllEnvs(); +}); + +describe("provider credential table", () => { + it("covers every registry provider with env var names or an ambient note", () => { + const registry = createPiModelRegistry(authStorage()); + const providers = registry.listProviders(); + expect(providers.length).toBeGreaterThan(30); + const uncovered = providers.filter((provider) => ( + getPiApiKeyEnvVarNames(provider).length === 0 && getAmbientCredentialNote(provider) === undefined + )); + expect(uncovered).toEqual([]); + }); + + it("maps the previously missing providers and keeps the bearer-token exclusion", () => { + expect(getPiApiKeyEnvVarNames("qwen-token-plan")).toEqual(["QWEN_TOKEN_PLAN_API_KEY"]); + expect(getPiApiKeyEnvVarNames("qwen-token-plan-cn")).toEqual(["QWEN_TOKEN_PLAN_CN_API_KEY"]); + expect(getPiApiKeyEnvVarNames("qwen-token-plan-individual")).toEqual(["QWEN_TOKEN_PLAN_API_KEY"]); + expect(getPiApiKeyEnvVarNames("anthropic")).not.toContain("ANTHROPIC_AUTH_TOKEN"); + expect(getPiCredentialEnvVarNames("anthropic")).toEqual(["ANTHROPIC_AUTH_TOKEN", "ANTHROPIC_OAUTH_TOKEN", "ANTHROPIC_API_KEY"]); + }); + + it("describes how to authenticate each kind of provider", () => { + expect(describeProviderCredentials("google")).toBe("set GEMINI_API_KEY (or run `codegenie provider login google`)"); + expect(describeProviderCredentials("amazon-bedrock")).toContain("AWS_ACCESS_KEY_ID"); + expect(describeProviderCredentials("google-vertex")).toContain("GOOGLE_CLOUD_API_KEY or Application Default Credentials"); + expect(describeProviderCredentials("openai-codex")).toContain("self-hosted runner"); + }); +}); + +describe("model resolution failure reasons", () => { + it("distinguishes unknown, deprecated, and credential-less models", () => { + clearProviderEnv("anthropic"); + const adapter = createRealPiAiAdapter({ authStorage: authStorage() }); + expect(adapter.resolveModel({ provider: "anthropic", model: "claude-opus-5" })).toBeUndefined(); + + const missing = adapter.explainUnresolvedModel?.({ provider: "anthropic", model: "claude-opus-5" }); + expect(missing).toEqual({ kind: "missing_credentials", provider: "anthropic", model: "claude-opus-5" }); + expect(modelResolutionMessage(missing!)).toBe( + "no credentials for provider anthropic: set ANTHROPIC_API_KEY (or run `codegenie provider login anthropic`)" + ); + + const unknown = adapter.explainUnresolvedModel?.({ provider: "anthropic", model: "claude-opus-5:hgh" }); + expect(unknown).toEqual({ kind: "unknown_model", provider: "anthropic", model: "claude-opus-5:hgh" }); + expect(modelResolutionMessage(unknown!)).toContain("unknown model anthropic/claude-opus-5:hgh"); + + const deprecated = adapter.explainUnresolvedModel?.({ provider: "anthropic", model: "claude-3-opus-20240229" }); + expect(deprecated).toMatchObject({ kind: "deprecated_model" }); + expect(modelResolutionMessage(deprecated!)).toBe("model anthropic/claude-3-opus-20240229 is deprecated; choose a current model"); + }); + + it("does not report a nonexistent provider as missing credentials", () => { + const adapter = createRealPiAiAdapter({ authStorage: authStorage() }); + expect(adapter.explainUnresolvedModel?.({ provider: "opnerouter" })).toEqual({ kind: "unresolved", provider: "opnerouter" }); + expect(adapter.explainUnresolvedModel?.({ provider: "opnerouter", model: "x" })).toMatchObject({ kind: "unknown_model" }); + clearProviderEnv("anthropic"); + expect(adapter.explainUnresolvedModel?.({ provider: "anthropic" })).toEqual({ kind: "missing_credentials", provider: "anthropic" }); + }); + + it("never labels an unexpected lookup error as missing credentials", () => { + const throwing = { + getModel: () => { + throw new Error("registry exploded"); + }, + getModels: () => [], + getProviders: () => [], + getProvider: () => undefined + } as unknown as Models; + const adapter = createRealPiAiAdapter({ authStorage: authStorage(), models: throwing }); + const failure = adapter.explainUnresolvedModel?.({ provider: "anthropic", model: "claude-opus-5" }); + expect(failure).toMatchObject({ kind: "unresolved" }); + expect(modelResolutionMessage(failure!)).toContain("no usable LLM model could be resolved"); + }); + + it("resolves normally when credentials exist, leaving no failure to explain", () => { + clearProviderEnv("anthropic"); + vi.stubEnv("ANTHROPIC_API_KEY", "sk-ant-test-only"); + const adapter = createRealPiAiAdapter({ authStorage: authStorage() }); + expect(adapter.resolveModel({ provider: "anthropic", model: "claude-opus-5" })?.apiKey).toBe("sk-ant-test-only"); + }); + + it("throws config_error with the specific message and a modelResolution marker", () => { + function build(adapter: PiAiAdapter): void { + createPiRunner({ + llmConfig: { provider: "anthropic", model: "claude-opus-5", maxConcurrentCalls: 1 }, + telemetry: telemetry(), + logger: logger(), + runSignal: new AbortController().signal, + adapter, + hooks: { checkpoint: () => "ok", onUsage: vi.fn() } + }); + } + const base = { resolveModel: () => undefined, complete: vi.fn(), validateToolCall: vi.fn() }; + + let caught: unknown; + try { + build({ ...base, explainUnresolvedModel: () => ({ kind: "missing_credentials", provider: "anthropic", model: "claude-opus-5" }) }); + } catch (error) { + caught = error; + } + expect(caught).toBeInstanceOf(CodegenieError); + expect(caught).toMatchObject({ + code: "config_error", + message: expect.stringContaining("no credentials for provider anthropic: set ANTHROPIC_API_KEY"), + context: { modelResolution: "missing_credentials" } + }); + + // Adapters without an explanation, or whose explanation throws, keep the + // generic message — never a guessed "missing credentials". + expect(() => build(base)).toThrow(/no usable LLM model could be resolved/u); + expect(() => build({ ...base, explainUnresolvedModel: () => { throw new Error("boom"); } })) + .toThrow(/no usable LLM model could be resolved/u); + }); +}); + +describe("llm-api-key routing reaches codegenie's resolver", () => { + // A stored login is refused by the Action instead (pi-ai lets it own the + // provider ahead of env vars); see the github-action entrypoint tests. + it("wins over competing env vars", () => { + vi.stubEnv("ANTHROPIC_AUTH_TOKEN", "bearer-token-competing"); + vi.stubEnv("ANTHROPIC_OAUTH_TOKEN", "oauth-token-competing"); + vi.stubEnv("ANTHROPIC_API_KEY", "stale-native-key"); + vi.stubEnv("LLM_API_KEY", "explicit-llm-api-key"); + const config = resolveModelConfig("opus", parseModelAliases("opus: anthropic/claude-opus-5\nsonnet: anthropic/claude-sonnet-5")); + + expect(applyLlmApiKey(process.env, config)).toBe("anthropic"); + + expect(process.env.ANTHROPIC_AUTH_TOKEN).toBeUndefined(); + expect(process.env.ANTHROPIC_OAUTH_TOKEN).toBeUndefined(); + expect(process.env.ANTHROPIC_API_KEY).toBe("explicit-llm-api-key"); + const model = createRealPiAiAdapter({ authStorage: authStorage() }).resolveModel({ provider: "anthropic", model: "claude-opus-5" }); + expect(model?.apiKey).toBe("explicit-llm-api-key"); + }); +}); + +function telemetry(): TelemetryRecorder { + return { + runId: "model-resolution", + runDir: undefined, + event: vi.fn(), + recordModelCall: vi.fn(), + recordToolCall: vi.fn(() => "tc-1"), + writeArtifact: vi.fn(async () => undefined), + writeDebug: vi.fn(async () => undefined), + flush: vi.fn(async () => undefined) + } as unknown as TelemetryRecorder; +} + +function logger(): Logger { + const sink = vi.fn(); + return { debug: sink, info: sink, warn: sink, error: sink }; +} From aaa6a230a5339a8ddf4a1964322a68dfdb11d2d3 Mon Sep 17 00:00:00 2001 From: Peter Kieltyka Date: Fri, 25 Sep 2026 12:26:29 -0400 Subject: [PATCH 2/3] fix(action): address PR #37 review of model aliases - llm-api-key's single-key rule now requires one API-key env var rather than one provider id, so provider pairs that share a var (moonshotai / moonshotai-cn, opencode / opencode-go, cloudflare-*) work with one key; competing vars are cleared and the stored-login guard runs per provider. - Validate model/models after the trigger gate, so a broken block fails real triggers but unrelated comments skip instead of going red. - Restore discriminated decision-record variants. - Reuse providerKnown for alias provider validation. - Normalize the requested alias only in the event gate, stripping surrounding quotes/backticks/brackets and trailing punctuation; alias names must start and end with a letter or digit so all stay reachable. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01NxBabX2aBd6eRLJryVci78 --- README.md | 2 +- ...-model-aliases-and-provider-credentials.md | 16 +++++ src/github-action/entrypoint.ts | 66 ++++++++++++------- src/github-action/event-gate.ts | 9 ++- src/github-action/models.ts | 51 +++++++------- src/provider/provider-services.ts | 2 +- tests/github-action.test.ts | 47 ++++++++++--- tests/model-resolution.test.ts | 2 +- 8 files changed, 131 insertions(+), 64 deletions(-) diff --git a/README.md b/README.md index d272ece..60ea7a6 100644 --- a/README.md +++ b/README.md @@ -140,7 +140,7 @@ List named models once. `model` stays the default for automatic reviews and a ba opus: anthropic/claude-opus-5:high ``` -- **Grammar.** The first word after the trigger phrase, on the same line, is looked up in `models` (case-insensitive). Nothing else in the comment is read. Comments cannot supply a model spec, a reasoning level or any other option; to offer a reasoning variant, add an alias for it (`opus-max: anthropic/claude-opus-5:max`). Without `models`, text after the trigger phrase is ignored exactly as before. +- **Grammar.** The first word after the trigger phrase, on the same line, is looked up in `models` (case-insensitive; surrounding quotes or backticks and trailing punctuation are ignored, so `` `opus` `` and `opus.` both work). Nothing else in the comment is read. Comments cannot supply a model spec, a reasoning level or any other option; to offer a reasoning variant, add an alias for it (`opus-max: anthropic/claude-opus-5:max`). Without `models`, text after the trigger phrase is ignored exactly as before. - **Unknown names.** `codegenie review opsu` runs no review. codegenie replies with the configured names (`Unknown model. Available: luna (default), deepseek, …`), and only to collaborators who pass the permission check. - **Keys.** `models` holds model specs only — never put keys in it. With several providers, set each provider's env var on the step, as above. If every configured model uses one provider, a single `llm-api-key` covers them all. When `llm-api-key` is set and the models span several providers, every run fails with a configuration error instead of sending one provider's key to another. On a self-hosted runner where someone ran `codegenie provider login` for that provider, the stored login would take precedence, so the run fails and says to log out there or drop `llm-api-key`. - **Precedence change.** `llm-api-key` now wins over a provider env var for the same provider. Older releases preferred the env var. A workflow that sets both, with different values, now uses `llm-api-key`. diff --git a/specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md b/specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md index 59c0f17..0b41c97 100644 --- a/specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md +++ b/specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md @@ -269,3 +269,19 @@ Likely files: `src/github-action/event-gate.ts`, `src/github-action/entrypoint.t - **Mutation checks:** removing the failure-body sanitizer, or the credential clearing, fails the new tests. Validation: `pnpm test` passed **1,583 tests across 69 files**, including actionlint on the merged example. `make evals` passed **39 synthetic tests**. Typecheck, build and `git diff --check` passed, and `models.md` regenerates identically. The built CLI printed the missing-credentials, unknown-model and deprecated-model messages on a real commit range, failing before any model call. No paid inference, pushes or GitHub posts were made. Dogfooding on a test PR (bare trigger, alias, unknown alias, alias with missing credentials) and the release remain with the owner. + +### Follow-up PR review (#37) + +A second high-effort review of the PR diff led to these changes (owner-approved; the stored-login guard is kept): + +- **One env var, not one provider.** The `llm-api-key` single-key rule now requires every selectable model to read the same API-key env var, not the same provider id. Provider pairs that share a var (`moonshotai`/`moonshotai-cn`, `opencode`/`opencode-go`, the two `cloudflare-*` providers) work with one key. Competing vars are cleared, and the stored-login guard runs, for every routed provider. +- **Validate after the trigger gate.** `model`/`models` are validated after the trigger gate instead of at flag parsing. A broken block still fails every real trigger, including every push, but unrelated comments ("LGTM") skip instead of failing. +- **Record types restored.** The decision-record union is back to discriminated variants (authorized, denied, unknown alias), so impossible records don't type-check. +- **Shared provider check.** Alias validation reuses `providerKnown` from `provider-services.ts`, which is now exported, instead of its own copy. +- **Token normalization in one place.** The event gate is the only place the alias token is normalized. It lowercases the token and strips surrounding quotes, backticks and brackets plus trailing punctuation, so `` `opus` `` and `opus.` select `opus`. To keep every alias reachable after that stripping, alias names must now start and end with a letter or digit. +- **Unchanged, as by-design or owner decisions:** + - Other words after the phrase get the reply when `models` is set. + - Pre-claim configuration errors reach only the job log, as an invalid `model` always did. + - The stale "Reviewing…" comment left by any superseding comment is the plan 97 residual. + - The optional `explainUnresolvedModel` re-resolution is kept. + - An Action-only explicit-key alternative to the stored-login guard exists, but it needs a new path for the key into the review; the owner kept the guard. diff --git a/src/github-action/entrypoint.ts b/src/github-action/entrypoint.ts index c709568..b423cc3 100644 --- a/src/github-action/entrypoint.ts +++ b/src/github-action/entrypoint.ts @@ -71,7 +71,10 @@ type GitHubActionInputs = { postInlineComments: boolean; preflightOnly: boolean; botLogin?: string; - models: ModelConfig; + // Raw `model` / `models` inputs, validated only after the trigger gate so + // a bad block cannot fail unrelated comment events. + modelInput?: string; + modelsInput: string; reviewPassthrough: string[]; }; @@ -115,6 +118,10 @@ export async function executeGitHubActionCommand( return; } + // A configuration error still fails every real trigger (including every + // push), but "LGTM" comments skip above instead of going red. + const models: ModelConfig = resolveModelConfig(inputs.modelInput, parseModelAliases(inputs.modelsInput)); + const comments = opts.issueComments ?? createIssueCommentClient(repoRoot, repoFullName); // Payload association fields are attacker-visible history; the live @@ -155,9 +162,9 @@ export async function executeGitHubActionCommand( // Resolved only after authorization, so unauthorized commenters never get // a reply. An unlisted alias gets fixed text from the workflow's own list. - const selected = selectModel(inputs.models, decision.requestedAlias); + const selected = selectModel(models, decision.requestedAlias); if (selected.kind === "unknown_alias") { - await comments.createComment(decision.prNumber, renderUnknownAliasReply(inputs.models)); + await comments.createComment(decision.prNumber, renderUnknownAliasReply(models)); const reason = "unknown model alias"; write(`github-action: skipped — ${reason}\n`); writeDecisionRecord(write, { ...decisionFields, run: false, reason }); @@ -172,15 +179,16 @@ export async function executeGitHubActionCommand( return; } - const keyProvider = applyLlmApiKey(env, inputs.models); - if (keyProvider !== undefined) { + const keyProviders = applyLlmApiKey(env, models); + if (keyProviders.length > 0) { // pi-ai lets a stored login own its provider ahead of env vars, which // would silently bypass llm-api-key. Only self-hosted runners can have one. const storage = opts.authStorage ?? createFileAuthStorage(getCodegeniePaths(undefined, env)); - if (storage.get(keyProvider) !== undefined) { + const overriding = keyProviders.find((provider) => storage.get(provider) !== undefined); + if (overriding !== undefined) { throw new CodegenieError( "invalid_args", - `a stored codegenie login for ${keyProvider} on this runner would override llm-api-key; run \`codegenie provider logout ${keyProvider}\` on the runner, or unset llm-api-key to use the stored login` + `a stored codegenie login for ${overriding} on this runner would override llm-api-key; run \`codegenie provider logout ${overriding}\` on the runner, or unset llm-api-key to use the stored login` ); } } @@ -312,11 +320,9 @@ export function parseGitHubActionArgs(argv: string[]): GitHubActionInputs { allowedUsers: [], postInlineComments: true, preflightOnly: false, - models: { aliases: new Map() }, + modelsInput: "", reviewPassthrough: [] }; - let modelInput: string | undefined; - let modelsInput = ""; const passthroughFlags = new Set(["--depth", "--lens", "--max-time", "--budget-boost"]); for (let index = 0; index < argv.length; index += 2) { @@ -338,9 +344,9 @@ export function parseGitHubActionArgs(argv: string[]): GitHubActionInputs { } else if (flag === "--preflight-only") { inputs.preflightOnly = parseBoolean(flag, value); } else if (flag === "--model") { - modelInput = value; + inputs.modelInput = value; } else if (flag === "--models") { - modelsInput = value; + inputs.modelsInput = value; } else if (flag === "--bot-login") { if (value.trim() !== "") { inputs.botLogin = value.trim(); @@ -356,8 +362,6 @@ export function parseGitHubActionArgs(argv: string[]): GitHubActionInputs { if (inputs.triggerPhrase.trim() === "") { throw new CodegenieError("invalid_args", "--trigger-phrase must not be empty"); } - // `model` and `models` are validated together, after every flag is read. - inputs.models = resolveModelConfig(modelInput, parseModelAliases(modelsInput)); return inputs; } @@ -562,21 +566,35 @@ function writeFailureFile(filePath: string | undefined, contents: string): void } } +type AuthorizedRecordFields = { + eventName: string; + lane: AuthorizedDecision["lane"]; + prNumber: number; + actor: string; + association: string; +}; + type DecisionRecord = | { eventName: string; run: false; reason: string } - | { - eventName: string; - run: boolean; - reason?: string; - lane: AuthorizedDecision["lane"]; - prNumber: number; - actor: string; - association: string; + | (AuthorizedRecordFields & { + run: true; actorAllowlisted: boolean; - permissionCheck: PermissionCheck | "denied"; + permissionCheck: PermissionCheck; modelAlias?: string; modelSpec?: string; - }; + }) + | (AuthorizedRecordFields & { + run: false; + reason: string; + actorAllowlisted: false; + permissionCheck: "denied"; + }) + | (AuthorizedRecordFields & { + run: false; + reason: "unknown model alias"; + actorAllowlisted: boolean; + permissionCheck: PermissionCheck; + }); function modelRecordFields(selection: ModelSelection | undefined): { modelAlias?: string; modelSpec?: string } { return { diff --git a/src/github-action/event-gate.ts b/src/github-action/event-gate.ts index 63c9ed4..2e00ef2 100644 --- a/src/github-action/event-gate.ts +++ b/src/github-action/event-gate.ts @@ -58,14 +58,17 @@ export function decideTrigger(eventName: string, payload: unknown, rules: Trigge } // The first whitespace-delimited token on the trigger phrase's own line, -// lowercased. Later lines never count, so a comment that continues on the -// next line still means "default". +// lowercased, with surrounding quotes/backticks/brackets and trailing +// punctuation stripped ("`opus`", "opus." → "opus"). This is the only place the +// token is normalized. Later lines never count, so a comment that continues on +// the next line still means "default". export function requestedAliasFromComment(body: string, phrase: string): string | undefined { if (!matchesTriggerPhrase(body, phrase)) { return undefined; } const rest = body.trim().slice(phrase.trim().length); - const token = (rest.split(/\r?\n/u)[0] ?? "").trim().split(/\s+/u)[0] ?? ""; + const token = ((rest.split(/\r?\n/u)[0] ?? "").trim().split(/\s+/u)[0] ?? "") + .replace(/^[`'"([]+|[`'")\].,;:!?]+$/gu, ""); return token === "" ? undefined : token.toLowerCase(); } diff --git a/src/github-action/models.ts b/src/github-action/models.ts index 1d4b17d..95952b6 100644 --- a/src/github-action/models.ts +++ b/src/github-action/models.ts @@ -3,11 +3,8 @@ // only ever selects an alias the workflow author listed — it never becomes a // model spec, reasoning level, or flag. import { isMap, isScalar, LineCounter, parseDocument } from "yaml"; -import { - getCodegeniePiModels, - getPiApiKeyEnvVarName, - getPiCredentialEnvVarNames -} from "../provider/pi-ai-models.js"; +import { getPiApiKeyEnvVarName, getPiCredentialEnvVarNames } from "../provider/pi-ai-models.js"; +import { providerKnown } from "../provider/provider-services.js"; import { splitReasoningSuffix } from "../provider/reasoning.js"; import { CodegenieError } from "../util/errors.js"; @@ -29,7 +26,9 @@ export type ModelConfig = { defaultModel?: ModelSelection; }; -const ALIAS_NAME_PATTERN = /^[A-Za-z0-9][A-Za-z0-9._-]{0,31}$/u; +// Starts and ends with a letter or digit, so the comment-side stripping of +// surrounding punctuation (event-gate) can never make an alias unreachable. +const ALIAS_NAME_PATTERN = /^[A-Za-z0-9](?:[A-Za-z0-9._-]{0,30}[A-Za-z0-9])?$/u; // One model spec instead of separate provider/model/reasoning inputs: // `provider/model[:reasoning]`, e.g. `anthropic/claude-opus-5:xhigh`. @@ -60,7 +59,7 @@ export function formatModelSpec(spec: ModelSpec): string { // masks secret values in the log regardless. export function parseModelAliases( text: string, - providerExists: (provider: string) => boolean = defaultProviderExists + providerExists: (provider: string) => boolean = providerKnown ): Map { const aliases = new Map(); if (text.trim() === "") { @@ -132,6 +131,7 @@ export function resolveModelConfig(modelInput: string | undefined, aliases: Map< // The model for this event. Without aliases, comment text is ignored exactly // as before; with aliases, an unlisted token is "unknown" (the caller replies). +// `requestedAlias` arrives normalized by the event gate. export function selectModel( config: ModelConfig, requestedAlias: string | undefined @@ -139,8 +139,8 @@ export function selectModel( if (config.aliases.size === 0 || requestedAlias === undefined) { return { kind: "selected", ...(config.defaultModel !== undefined ? { selection: config.defaultModel } : {}) }; } - const spec = config.aliases.get(requestedAlias.toLowerCase()); - return spec === undefined ? { kind: "unknown_alias" } : { kind: "selected", selection: { alias: requestedAlias.toLowerCase(), spec } }; + const spec = config.aliases.get(requestedAlias); + return spec === undefined ? { kind: "unknown_alias" } : { kind: "selected", selection: { alias: requestedAlias, spec } }; } // Fixed text built only from configured alias names (restricted charset), so @@ -153,14 +153,15 @@ export function renderUnknownAliasReply(config: ModelConfig): string { } // `llm-api-key` (LLM_API_KEY) is authoritative when set: every selectable -// model must share one API-key provider, that provider's competing credential -// vars are cleared, and the key is written unconditionally. Unset, provider -// credentials resolve exactly as they always have. Returns the provider the -// key was routed to, if any. -export function applyLlmApiKey(env: NodeJS.ProcessEnv, config: ModelConfig): string | undefined { +// model must read its key from one env var (usually one provider; a few +// provider pairs share a var, e.g. moonshotai/moonshotai-cn), those providers' +// competing credential vars are cleared, and the key is written +// unconditionally. Unset, provider credentials resolve exactly as they always +// have. Returns the providers the key was routed to (empty when unset). +export function applyLlmApiKey(env: NodeJS.ProcessEnv, config: ModelConfig): string[] { const key = env.LLM_API_KEY; if (key === undefined || key === "") { - return undefined; + return []; } const selectable = [ ...(config.defaultModel !== undefined ? [config.defaultModel.spec] : []), @@ -173,29 +174,27 @@ export function applyLlmApiKey(env: NodeJS.ProcessEnv, config: ModelConfig): str ); } const providers = [...new Set(selectable.map((spec) => spec.provider as string))]; - if (providers.length > 1) { + const envVarNames = new Set(providers.map((provider) => getPiApiKeyEnvVarName(provider))); + if (envVarNames.size > 1) { throw new CodegenieError( "invalid_args", `llm-api-key is a single key, but models use providers ${providers.join(", ")}. Remove llm-api-key and set each provider's env var (see models.md#credentials).` ); } - const provider = providers[0] as string; - const envVarName = getPiApiKeyEnvVarName(provider); + const [envVarName] = envVarNames; if (envVarName === undefined) { throw new CodegenieError( "invalid_args", - `provider ${provider} does not accept an API key; set its native credentials instead of LLM_API_KEY` + `provider ${providers.join(", ")} does not accept an API key; set its native credentials instead of LLM_API_KEY` ); } - for (const name of getPiCredentialEnvVarNames(provider)) { - delete env[name]; + for (const provider of providers) { + for (const name of getPiCredentialEnvVarNames(provider)) { + delete env[name]; + } } env[envVarName] = key; - return provider; -} - -function defaultProviderExists(provider: string): boolean { - return getCodegeniePiModels().getProvider(provider) !== undefined; + return providers; } function modelsError(detail: string): CodegenieError { diff --git a/src/provider/provider-services.ts b/src/provider/provider-services.ts index 1c80aab..407994c 100644 --- a/src/provider/provider-services.ts +++ b/src/provider/provider-services.ts @@ -975,7 +975,7 @@ function resolveProviderAlias(provider: string): string { return PROVIDER_ALIASES[provider] ?? provider; } -function providerKnown(provider: string, models: Pick = getCodegeniePiModels()): boolean { +export function providerKnown(provider: string, models: Pick = getCodegeniePiModels()): boolean { return models.getProvider(provider) !== undefined; } diff --git a/tests/github-action.test.ts b/tests/github-action.test.ts index 2423348..70eefd7 100644 --- a/tests/github-action.test.ts +++ b/tests/github-action.test.ts @@ -173,6 +173,13 @@ describe("github-action event gate", () => { expect(requestedAliasFromComment("codegenie review \r\nopus", "codegenie review")).toBeUndefined(); expect(requestedAliasFromComment("codegenie reviewopus", "codegenie review")).toBeUndefined(); expect(requestedAliasFromComment("please codegenie review opus", "codegenie review")).toBeUndefined(); + // Surrounding quotes/backticks/brackets and trailing punctuation are stripped. + expect(requestedAliasFromComment("codegenie review `opus`", "codegenie review")).toBe("opus"); + expect(requestedAliasFromComment("codegenie review opus.", "codegenie review")).toBe("opus"); + expect(requestedAliasFromComment("codegenie review \"GLM\"!", "codegenie review")).toBe("glm"); + expect(requestedAliasFromComment("codegenie review (deep-seek),", "codegenie review")).toBe("deep-seek"); + expect(requestedAliasFromComment("codegenie review gpt-5.5", "codegenie review")).toBe("gpt-5.5"); + expect(requestedAliasFromComment("codegenie review ...", "codegenie review")).toBeUndefined(); expect(decideTrigger("issue_comment", issueCommentPayload({ body: "codegenie review GLM" }), RULES)).toMatchObject({ run: true, @@ -220,6 +227,8 @@ describe("github-action model aliases", () => { const cases: Array<[string, RegExp]> = [ ["Opus: anthropic/claude-opus-5\nopus: anthropic/claude-sonnet-5", /line 2: duplicate alias opus/u], ["-bad: anthropic/claude-opus-5", /line 1: alias names must match/u], + ["bad.: anthropic/claude-opus-5", /line 1: alias names must match/u], + ["bad-: anthropic/claude-opus-5", /line 1: alias names must match/u], ["has space: anthropic/claude-opus-5", /line 1: alias names must match/u], ["opus:\n model: anthropic/claude-opus-5", /line 1: alias opus must map to a provider\/model/u], ["opus: [anthropic/claude-opus-5]", /line 1: alias opus must map to a provider\/model/u], @@ -261,7 +270,7 @@ describe("github-action model aliases", () => { it("selects the requested alias, ignores tokens without models, and lists names for unknowns", () => { const config = resolveModelConfig("luna", parseModelAliases(MODELS)); expect(selectModel(config, undefined)).toEqual({ kind: "selected", selection: config.defaultModel }); - expect(selectModel(config, "OPUS")).toMatchObject({ kind: "selected", selection: { alias: "opus" } }); + expect(selectModel(config, "opus")).toMatchObject({ kind: "selected", selection: { alias: "opus" } }); expect(selectModel(config, "opsu")).toEqual({ kind: "unknown_alias" }); const legacy = resolveModelConfig("openrouter/deepseek/deepseek-v4.1-flash:max", new Map()); expect(selectModel(legacy, "opsu")).toEqual({ kind: "selected", selection: legacy.defaultModel }); @@ -273,14 +282,15 @@ describe("github-action model aliases", () => { ); }); - it("validates model and models together when parsing action flags", () => { - expect(() => parseGitHubActionArgs(["--models", MODELS])).toThrow(/--models requires --model/u); - expect(() => parseGitHubActionArgs(["--models", MODELS, "--model", "sonnet"])).toThrow(/not one of the configured/u); - const parsed = parseGitHubActionArgs(["--models", MODELS, "--model", "luna"]); - expect(parsed.models.defaultModel?.alias).toBe("luna"); - expect(parsed.models.aliases.size).toBe(3); + it("keeps model and models raw at flag parsing; they are validated after the trigger gate", () => { + // Validation happens in the entrypoint (see "validates model and models + // only for real triggers"), so parsing never throws on these values. + expect(parseGitHubActionArgs(["--models", "opus: [broken", "--model", "sonnet"])).toMatchObject({ + modelInput: "sonnet", + modelsInput: "opus: [broken" + }); // The composite action always forwards both flags, possibly empty. - expect(parseGitHubActionArgs(["--model", "", "--models", ""]).models).toEqual({ aliases: new Map() }); + expect(parseGitHubActionArgs(["--model", "", "--models", ""])).toMatchObject({ modelInput: "", modelsInput: "" }); }); }); @@ -1264,6 +1274,15 @@ describe("github-action entrypoint", () => { }); expect(String((caught as Error).message)).not.toContain(secret); + // Different providers that read the same env var share one key. + const sharedVar = resolveModelConfig("kimi", parseModelAliases([ + "kimi: moonshotai/kimi-k2.5", + "kimi-cn: moonshotai-cn/kimi-k2.5" + ].join("\n"))); + const shared: NodeJS.ProcessEnv = { LLM_API_KEY: secret }; + expect(applyLlmApiKey(shared, sharedVar)).toEqual(["moonshotai", "moonshotai-cn"]); + expect(shared.MOONSHOT_API_KEY).toBe(secret); + const sameProvider = resolveModelConfig("luna", parseModelAliases([ "luna: openrouter/openai/gpt-6-luna:xhigh", "glm: openrouter/z-ai/glm-5.3:max" @@ -1334,6 +1353,18 @@ describe("github-action entrypoint", () => { expect(opus.output).toContain('"modelAlias":"opus"'); }); + it("validates model and models only for real triggers", async () => { + const broken = ["--model", "luna", "--models", "luna: [openrouter/openai/gpt-6-luna"]; + const unrelated = await runWithModels(issueCommentPayload({ body: "LGTM, thanks!" }), "issue_comment", broken); + expect(unrelated.output).toContain("skipped — comment does not match the trigger phrase"); + expect(unrelated.calls).toHaveLength(0); + + await expect(runWithModels(issueCommentPayload(), "issue_comment", broken)).rejects.toThrow(/--models invalid YAML at line 1/u); + await expect(runWithModels(pullRequestPayload(), "pull_request", ["--models", MODELS, "--model", "sonnet"])) + .rejects.toThrow(/not one of the configured model aliases/u); + await expect(runWithModels(pullRequestPayload(), "pull_request", ["--models", MODELS])).rejects.toThrow(/--models requires --model/u); + }); + it("replies once to an unknown alias without claiming the status comment or reviewing", async () => { const result = await runWithModels(issueCommentPayload({ body: "codegenie review opsu @someone" }), "issue_comment", MODEL_ARGS); expect(result.argv).toBeUndefined(); diff --git a/tests/model-resolution.test.ts b/tests/model-resolution.test.ts index 2bba1bf..d325c01 100644 --- a/tests/model-resolution.test.ts +++ b/tests/model-resolution.test.ts @@ -158,7 +158,7 @@ describe("llm-api-key routing reaches codegenie's resolver", () => { vi.stubEnv("LLM_API_KEY", "explicit-llm-api-key"); const config = resolveModelConfig("opus", parseModelAliases("opus: anthropic/claude-opus-5\nsonnet: anthropic/claude-sonnet-5")); - expect(applyLlmApiKey(process.env, config)).toBe("anthropic"); + expect(applyLlmApiKey(process.env, config)).toEqual(["anthropic"]); expect(process.env.ANTHROPIC_AUTH_TOKEN).toBeUndefined(); expect(process.env.ANTHROPIC_OAUTH_TOKEN).toBeUndefined(); From 19505c06132c87d142a039eafc6f686f0c022fdc Mon Sep 17 00:00:00 2001 From: Peter Kieltyka Date: Fri, 25 Sep 2026 12:30:17 -0400 Subject: [PATCH 3/3] 0.7.0 Bump the package to 0.7.0 and pin the documented action references to v0.7.0, the first release with the `models` input. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01NxBabX2aBd6eRLJryVci78 --- README.md | 4 ++-- examples/workflows/codegenie-review.yml | 2 +- package.json | 2 +- ...issue-125-github-model-aliases-and-provider-credentials.md | 2 +- 4 files changed, 5 insertions(+), 5 deletions(-) diff --git a/README.md b/README.md index 60ea7a6..3f8cd7c 100644 --- a/README.md +++ b/README.md @@ -105,7 +105,7 @@ jobs: with: ref: ${{ github.event.pull_request.base.sha }} # trusted base; PR head is fetched as review data fetch-depth: 0 - - uses: 0xPolygon/codegenie@v0.6.3 + - uses: 0xPolygon/codegenie@v0.7.0 with: # Works with any model! model: "openrouter/openai/gpt-6-luna:xhigh" @@ -126,7 +126,7 @@ The `model` input is one spec: `provider/model[:reasoning]` — any model in [mo List named models once. `model` stays the default for automatic reviews and a bare `codegenie review`. A collaborator can comment `codegenie review opus` to run that one review with the `opus` entry instead. ```yaml - - uses: 0xPolygon/codegenie@v0.6.3 + - uses: 0xPolygon/codegenie@v0.7.0 env: # one credential env var per provider in the list (see Credentials below) OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }} diff --git a/examples/workflows/codegenie-review.yml b/examples/workflows/codegenie-review.yml index 63cc4fc..65a6acf 100644 --- a/examples/workflows/codegenie-review.yml +++ b/examples/workflows/codegenie-review.yml @@ -45,7 +45,7 @@ jobs: ref: ${{ github.event.pull_request.base.sha || '' }} fetch-depth: 0 - - uses: 0xPolygon/codegenie@v0.6.3 + - uses: 0xPolygon/codegenie@v0.7.0 # One credential env var per provider used below; the name for every # provider is listed in models.md#credentials. env: diff --git a/package.json b/package.json index 62518ee..626a6ac 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "@0xsequence/codegenie", - "version": "0.6.3", + "version": "0.7.0", "description": "High-signal AI code review agent", "type": "module", "bin": { diff --git a/specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md b/specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md index 0b41c97..5798fbd 100644 --- a/specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md +++ b/specs/plans/125-issue-125-github-model-aliases-and-provider-credentials.md @@ -264,7 +264,7 @@ Likely files: `src/github-action/event-gate.ts`, `src/github-action/entrypoint.t - Found that pi-ai's stored-credential-first auth bypasses `llm-api-key` on self-hosted runners with a stored login. Resolved by the owner-chosen Action guard. - Found that a nonexistent provider was reported as "missing credentials". Fixed: it now gets the generic message. - Found that numeric alias names were rejected. Fixed with the failsafe schema. - - Not changed: the example and README pins stay `@v0.6.3`, which has no `models` input. The existing contract test forces them to the package version at release. The optional adapter method (rather than a reason-returning `resolveModel`) was kept to avoid touching the ~40 test adapters. + - Pins: the example and README pins stayed `@v0.6.3` (no `models` input) until the package was bumped to 0.7.0 on this branch; they now pin `@v0.7.0`, as the existing contract test requires. The optional adapter method (rather than a reason-returning `resolveModel`) was kept to avoid touching the ~40 test adapters. - Sound: the trust boundary, secrets on every published surface, backward compatibility and error labeling on the Action path. - **Mutation checks:** removing the failure-body sanitizer, or the credential clearing, fails the new tests.