Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions docs/llmservice/models/deepseek-v4-flash.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ import ActivityCard from '@site/src/components/ActivityCard';

## Overview

DeepSeek-V4-Flash is DeepSeek's high-efficiency open-source language model, released alongside V4-Pro on April 24, 2026 under the MIT License. With 284 billion total parameters and only 13 billion active parameters, it delivers performance within striking distance of V4-Pro at roughly 3.1x lower cost, making it one of the most cost-effective models available.
DeepSeek-V4-Flash is DeepSeek's high-efficiency open-source language model, released alongside V4-Pro on April 24, 2026 under the MIT License. With 284 billion total parameters and only 13 billion active parameters, it delivers performance within striking distance of V4-Pro at roughly one-third of the standard input and output price, making it one of the most cost-effective models available.

<ActivityCard
variant="free"
Expand All @@ -24,12 +24,12 @@ After the offer ends, the model will return to standard pricing. Offer end time,

* **Ultra-Efficient Architecture**: 284B total parameters with just 13B activated per forward pass, resulting in a compact 160GB download that runs on significantly less hardware than frontier models while maintaining strong performance.
* **1M-Token Context Window**: Shares the same 1-million-token context and 384K max output as V4-Pro, powered by the same CSA/HCA hybrid attention mechanism for efficient long-context inference.
* **Near-Pro Performance at Lower Cost**: Scores 79.0% on SWE-bench Verified, only 1.6 percentage points behind V4-Pro's 80.6%, while costing 0.28/0.56 Credits per input/output token.
* **Near-Pro Performance at Lower Cost**: Scores 79.0% on SWE-bench Verified, only 1.6 percentage points behind V4-Pro's 80.6%, while its standard reference price is 0.44/1.32 Credits per input/output token.
* **Flash-Max Reasoning Mode**: When given a larger thinking budget (384K+ context), V4-Flash-Max achieves comparable reasoning performance to V4-Pro, closing the gap on complex tasks.

## Best Use Cases

* **High-Volume API Workloads**: At 0.28 Credits per input token, Flash is ideal for applications that process large volumes of text where cost per query matters more than marginal accuracy gains.
* **High-Volume API Workloads**: With a standard reference input price of 0.44 Credits per token, Flash is ideal for applications that process large volumes of text where cost per query matters more than marginal accuracy gains.
* **Self-Hosted Deployments**: The 160GB model size and 13B active parameters make it feasible for on-premise or single-node GPU deployments, unlike larger frontier models.
* **Agentic Tool-Use Pipelines**: Strong tool-calling and coding capabilities paired with low latency make it well-suited for multi-step agent workflows where many LLM calls are chained together.

Expand All @@ -56,8 +56,8 @@ After the offer ends, the model will return to standard pricing. Offer end time,

| Model | Input (Credits/Token) | Cache Write (Credits/Token) | Cache Read (Credits/Token) | Output (Credits/Token) | Web Search (Credits/Use) | Billing Notes |
| :--- | --------------------: | --------------------------: | -------------------------: | ---------------------: | -----------------------: | :--- |
| **DeepSeek-V4-Flash** | `0.28` | `0.28` | `0.0056` | `0.56` | `-` | - |
| **DeepSeek-V4-Flash** | `0.44` | `0.44` | `0.0088` | `1.32` | `-` | Standard reference price; Cache Write: `1x` input; Cache Read: `0.02x` input |

:::info Pricing note
Prices shown in the documentation are B.AI standard reference prices for base billing purposes. B.AI may provide lower actual usage costs through top-up bonuses and account benefits. Specific prices, bonus Credits, and account benefits are subject to the platform display and final billing records.
The table shows the standard reference price for DeepSeek-V4-Flash. Its current limited-time offer applies `0 Credits` to all B.AI Chat and API usage. Time-based pricing is not yet enabled; final settlement prices and billing records are subject to the platform display. B.AI may provide lower actual usage costs through top-up bonuses and account benefits.
:::
4 changes: 2 additions & 2 deletions docs/llmservice/models/deepseek-v4-pro.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,8 +40,8 @@ DeepSeek-V4-Pro is DeepSeek's flagship open-source large language model, release

| Model | Input (Credits/Token) | Cache Write (Credits/Token) | Cache Read (Credits/Token) | Output (Credits/Token) | Web Search (Credits/Use) | Billing Notes |
| :--- | --------------------: | --------------------------: | -------------------------: | ---------------------: | -----------------------: | :--- |
| **DeepSeek V4 Pro** | `0.87` | `0.87` | `0.0087` | `1.74` | `-` | - |
| **DeepSeek V4 Pro** | `1.32` | `1.32` | `0.0132` | `3.96` | `-` | Currently billed at the Busy rate; Cache Write: `1x` input; Cache Read: `0.01x` input |

:::info Pricing note
Prices shown in the documentation are B.AI standard reference prices for base billing purposes. B.AI may provide lower actual usage costs through top-up bonuses and account benefits. Specific prices, bonus Credits, and account benefits are subject to the platform display and final billing records.
DeepSeek V4 Pro currently uses the Busy price. Idle pricing and peak/off-peak pricing are not yet enabled; when available, the platform will show the applicable pricing rules. Final settlement prices and billing records are subject to the platform display. B.AI may provide lower actual usage costs through top-up bonuses and account benefits.
:::
2 changes: 1 addition & 1 deletion docs/llmservice/models/gpt-5-5.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ GPT-5.5 is OpenAI's most capable model, released on April 23, 2026. Codenamed "S

### Known Limitations

* Significantly more expensive than open-source alternatives, roughly 6x the cost of DeepSeek V4 Pro for input tokens.
* Significantly more expensive than many open-source alternatives; direct cost comparisons should use the current model pricing.
* Lost the harder SWE-Bench Pro benchmark to Claude Opus 4.7 despite winning the standard SWE-Bench Verified headline.
* 2x input pricing for prompts exceeding 272K tokens increases costs substantially for long-context workloads.

Expand Down
6 changes: 3 additions & 3 deletions docs/llmservice/pricing-and-usage.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,8 +30,8 @@ The table below lists standard reference prices only. For current limited-time o
| GLM-5.2 | 1.40 | 1.40 | 0.28 | 4.40 | - |
| GLM-5.1 | 1.40 | 1.40 | 0.28 | 4.40 | - |
| DeepSeek V3.2 | 0.29 | 0.29 | 0.145 | 0.44 | - |
| DeepSeek-V4-Flash | 0.28 | 0.28 | 0.0056 | 0.56 | - |
| DeepSeek V4 Pro | 0.87 | 0.87 | 0.0087 | 1.74 | - |
| DeepSeek-V4-Flash | 0.44 | 0.44 | 0.0088 | 1.32 | - |
| DeepSeek V4 Pro | 1.32 | 1.32 | 0.0132 | 3.96 | - |
| Grok 4.6 | 2.00 | 2.00 | 0.50 | 6.00 | - |
| Grok 4.5 | 2.00 | 2.00 | 0.30 | 6.00 | - |
| GPT-5.6 Sol | 5.00 | 6.25 | 0.50 | 30.00 | 10,000 |
Expand Down Expand Up @@ -63,7 +63,7 @@ The table below lists standard reference prices only. For current limited-time o
| Gemini 3 Flash | 0.50 | 0.50 | 0.05 | 3.00 | 14,000 |

:::caution Main table scope
The main pricing table shows the currently effective standard reference price for each model. The `Cache Write` column represents the billing rate when cache writing occurs; it does not imply a unified cache TTL across all models. Cache behavior, retention time, long-context pricing, and extended caching options may vary by model provider. If a model has special caching rules, long-context pricing, 1-hour cache write pricing, or time-based pricing, please refer to the corresponding model detail page.
The main pricing table shows the currently effective standard reference price for each model. DeepSeek V4 Pro currently applies its Busy price; Idle pricing is not yet enabled. DeepSeek-V4-Flash is currently free on B.AI Chat and API under its limited-time offer, while its row shows the standard reference price. Time-based pricing is not yet enabled for either model. The `Cache Write` column represents the billing rate when cache writing occurs; it does not imply a unified cache TTL across all models. Cache behavior, retention time, long-context pricing, and extended caching options may vary by model provider. If a model has special caching rules, long-context pricing, 1-hour cache write pricing, or time-based pricing, please refer to the corresponding model detail page.
:::

:::info Pricing note
Expand Down
8 changes: 4 additions & 4 deletions docs/llmservice/promotions-and-pricing-notices.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,10 +53,10 @@ For a limited time, eligible requests are billed at 90% of the standard referenc
<ActivityCard
variant="adjustment"
title="DeepSeek API Pricing"
status="Pricing Adjustment Notice"
detail="Subject to Official Notice"
status="Current Pricing"
detail="V4 Pro: Busy Rate"
>
Due to a recent pricing adjustment by DeepSeek, B.AI plans to make a corresponding adjustment to pricing for DeepSeek API services. Please plan your usage accordingly.
DeepSeek-V4-Pro is currently billed at its Busy rate. Idle pricing and peak/off-peak pricing are not yet enabled.

The adjustment scope, effective date, and final prices are subject to the formal announcement and platform display.
DeepSeek-V4-Flash is currently free across B.AI Chat and API under its limited-time offer. Time-based pricing is not yet enabled. See the [DeepSeek-V4-Pro](./models/deepseek-v4-pro.md) and [DeepSeek-V4-Flash](./models/deepseek-v4-flash.md) model details for pricing information. Final billing is subject to the platform display.
</ActivityCard>
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ import ActivityCard from '@site/src/components/ActivityCard';

## 概述

DeepSeek-V4-Flash 是 DeepSeek 于 2026 年 4 月 24 日与 V4-Pro 同步发布的高效率开源大语言模型,采用 MIT License。该模型总参数量为 284B,但每次前向仅激活 13B 参数,以仅为 V4-Pro 约 1/3.1 的成本提供接近旗舰模型的性能,是当前极具性价比的模型之一。
DeepSeek-V4-Flash 是 DeepSeek 于 2026 年 4 月 24 日与 V4-Pro 同步发布的高效率开源大语言模型,采用 MIT License。该模型总参数量为 284B,但每次前向仅激活 13B 参数,以约为 V4-Pro 三分之一的标准输入和输出价格提供接近旗舰模型的性能,是当前极具性价比的模型之一。

<ActivityCard
variant="free"
Expand All @@ -24,12 +24,12 @@ DeepSeek-V4-Flash 是 DeepSeek 于 2026 年 4 月 24 日与 V4-Pro 同步发布

* **超高效率架构**:总参数量 284B,每次前向仅激活 13B 参数,模型下载体积约 160GB,相比前沿模型对硬件要求更低,同时保持出色性能。
* **100 万 Token 上下文窗口**:与 V4-Pro 一样支持 100 万上下文和 384K 最大输出,基于相同的 CSA/HCA 混合注意力机制,具备高效的长上下文推理能力。
* **接近 Pro 的性能与更低成本**:在 SWE-bench Verified 上达到 79.0%,仅比 V4-Pro 的 80.6% 低 1.6 个百分点,而输入/输出价格仅为 0.28 / 0.56 Credits。
* **接近 Pro 的性能与更低成本**:在 SWE-bench Verified 上达到 79.0%,仅比 V4-Pro 的 80.6% 低 1.6 个百分点,标准参考输入/输出价格为 0.44 / 1.32 Credits。
* **Flash-Max 推理模式**:在提供更大的思考预算(384K+ 上下文)时,V4-Flash-Max 可在复杂任务上逼近 V4-Pro 的推理能力。

## 适用场景

* **高并发 API 场景**:以每输入 token 仅 0.28 Credits 的成本,非常适合文本量大、对单次调用成本敏感的应用。
* **高并发 API 场景**:标准参考输入价格为每 token 0.44 Credits,适合文本量大、对单次调用成本敏感的应用。
* **自托管部署**:160GB 模型体积和 13B 激活参数使其更适合本地部署或单节点 GPU 场景,不像更大的前沿模型那样依赖重型基础设施。
* **Agent 工具调用链路**:强工具调用和编程能力,加上更低延迟,使其非常适合多步 Agent 工作流。

Expand All @@ -54,10 +54,10 @@ DeepSeek-V4-Flash 是 DeepSeek 于 2026 年 4 月 24 日与 V4-Pro 同步发布

## 积分消耗

| 模型名称 | 输入 (Credits/Token) | Cache Write (Credits/Token) | Cache Read (Credits/Token) | 输出 (Credits/Token) | 网页搜索(Credits/次) | 计费说明 |
| :--- | --------------------: | --------------------------: | -------------------------: | -------------------: | ---------------------: | :--- |
| **DeepSeek-V4-Flash** | `0.28` | `0.28` | `0.0056` | `0.56` | `-` | - |
| 模型名称 | 输入 (Credits/Token) | 缓存写入 (Credits/Token) | 缓存读取 (Credits/Token) | 输出 (Credits/Token) | 网页搜索(Credits/次) | 计费说明 |
| :------- | --------------------: | ------------------------: | ------------------------: | -------------------: | ---------------------: | :--- |
| **DeepSeek-V4-Flash** | `0.44` | `0.44` | `0.0088` | `1.32` | `-` | 标准参考价;缓存写入:输入价的 `1x`;缓存读取:输入价的 `0.02x` |

:::info 价格说明
文档价格为 B.AI 平台模型标准参考价,仅供基础计费说明使用。B.AI 可能会通过充值赠送及账户权益等方式,为用户提供更低的实际使用成本。具体价格、赠送积分及账户权益请以平台页面展示及最终账单为准
本表展示 DeepSeek-V4-Flash 的标准参考价。当前限时活动期间,B.AI Chat 和 API 的所有使用均按 `0 Credits` 结算。峰谷定价尚未启用;实际结算价格及最终账单以平台页面展示为准。B.AI 可能会通过充值赠送及账户权益等方式,为用户提供更低的实际使用成本。
:::
Original file line number Diff line number Diff line change
Expand Up @@ -38,10 +38,10 @@ DeepSeek-V4-Pro 是 DeepSeek 于 2026 年 4 月 24 日基于 MIT License 发布

## 积分消耗

| 模型名称 | 输入 (Credits/Token) | Cache Write (Credits/Token) | Cache Read (Credits/Token) | 输出 (Credits/Token) | 网页搜索(Credits/次) | 计费说明 |
| :--- | --------------------: | --------------------------: | -------------------------: | -------------------: | ---------------------: | :--- |
| **DeepSeek V4 Pro** | `0.87` | `0.87` | `0.0087` | `1.74` | `-` | - |
| 模型名称 | 输入 (Credits/Token) | 缓存写入 (Credits/Token) | 缓存读取 (Credits/Token) | 输出 (Credits/Token) | 网页搜索(Credits/次) | 计费说明 |
| :------- | --------------------: | ------------------------: | ------------------------: | -------------------: | ---------------------: | :--- |
| **DeepSeek V4 Pro** | `1.32` | `1.32` | `0.0132` | `3.96` | `-` | 当前按忙时价格结算;缓存写入:输入价的 `1x`;缓存读取:输入价的 `0.01x` |

:::info 价格说明
文档价格为 B.AI 平台模型标准参考价,仅供基础计费说明使用。B.AI 可能会通过充值赠送及账户权益等方式,为用户提供更低的实际使用成本。具体价格、赠送积分及账户权益请以平台页面展示及最终账单为准
DeepSeek V4 Pro 当前按忙时价格结算,闲时价格及峰谷定价尚未启用;相关能力上线后,平台将展示对应的计费规则。实际结算价格及最终账单以平台页面展示为准。B.AI 可能会通过充值赠送及账户权益等方式,为用户提供更低的实际使用成本。
:::
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ GPT-5.5 是 OpenAI 于 2026 年 4 月 23 日发布的最强模型,内部代号

### 已知限制

* 相比开源替代模型成本明显更高,输入价格约为 DeepSeek V4 Pro 的 6 倍
* 相比许多开源替代模型成本明显更高;直接成本比较应以模型当前价格为准
* 虽然在标准 SWE-Bench Verified 上领先,但在更难的 SWE-Bench Pro 基准上落后于 Claude Opus 4.7。
* 对于超过 272K token 的长上下文输入,输入价格会翻倍,显著提高整体成本。

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -30,8 +30,8 @@
| GLM-5.2 | 1.40 | 1.40 | 0.28 | 4.40 | - |
| GLM-5.1 | 1.40 | 1.40 | 0.28 | 4.40 | - |
| DeepSeek V3.2 | 0.29 | 0.29 | 0.145 | 0.44 | - |
| DeepSeek-V4-Flash | 0.28 | 0.28 | 0.0056 | 0.56 | - |
| DeepSeek V4 Pro | 0.87 | 0.87 | 0.0087 | 1.74 | - |
| DeepSeek-V4-Flash | 0.44 | 0.44 | 0.0088 | 1.32 | - |
| DeepSeek V4 Pro | 1.32 | 1.32 | 0.0132 | 3.96 | - |
| Grok 4.6 | 2.00 | 2.00 | 0.50 | 6.00 | - |
| Grok 4.5 | 2.00 | 2.00 | 0.30 | 6.00 | - |
| GPT-5.6 Sol | 5.00 | 6.25 | 0.50 | 30.00 | 10,000 |
Expand Down Expand Up @@ -63,7 +63,7 @@
| Gemini 3 Flash | 0.50 | 0.50 | 0.05 | 3.00 | 14,000 |

:::caution 价格总表说明
价格总表展示每个模型当前生效的标准参考价。表中的“缓存写入(Cache Write)”表示发生缓存写入时的计费价格,不代表所有模型使用统一的缓存有效期。不同模型厂商的缓存策略、缓存有效期、长上下文价格和扩展缓存能力可能不同;如模型存在特殊缓存规则、长上下文价格、1 小时缓存写入价格或分时间生效的价格,请以对应模型详情页说明为准。
价格总表展示每个模型当前生效的标准参考价。DeepSeek V4 Pro 当前按忙时价格结算,闲时价格尚未启用。DeepSeek-V4-Flash 当前在 B.AI Chat 和 API 中处于限时免费活动,表中展示其标准参考价。两款模型的峰谷定价尚未启用。表中的“缓存写入(Cache Write)”表示发生缓存写入时的计费价格,不代表所有模型使用统一的缓存有效期。不同模型厂商的缓存策略、缓存有效期、长上下文价格和扩展缓存能力可能不同;如模型存在特殊缓存规则、长上下文价格、1 小时缓存写入价格或分时间生效的价格,请以对应模型详情页说明为准。
:::

:::info 价格说明
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -53,10 +53,10 @@ import ActivityCard from '@site/src/components/ActivityCard';
<ActivityCard
variant="adjustment"
title="DeepSeek API 定价"
status="定价调整预告"
detail="以正式通知为准"
status="当前计费说明"
detail="V4 Pro:当前按忙时计费"
>
由于 DeepSeek 官方近期调整服务定价,B.AI 计划相应上调 DeepSeek API 服务价格。请根据业务需求合理安排使用
DeepSeek-V4-Pro 当前按忙时价格结算,闲时价格及峰谷定价尚未启用

具体调整范围、生效时间及最终价格以正式通知和平台页面展示为准
DeepSeek-V4-Flash 当前在限时免费活动期间,B.AI Chat 和 API 使用均免费;峰谷定价尚未启用。完整价格请查看 [DeepSeek-V4-Pro](./models/deepseek-v4-pro.md) 和 [DeepSeek-V4-Flash](./models/deepseek-v4-flash.md) 模型详情。最终账单以平台页面展示为准
</ActivityCard>
2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@x402-tron/docs",
"version": "1.3.21",
"version": "1.3.22",
"description": "x402-tron documentation",
"license": "MIT",
"resolutions": {
Expand Down
Loading