You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 155146f
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: en/ai-sre/im.mdx
+25-7Lines changed: 25 additions & 7 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -32,12 +32,14 @@ AI SRE's IM integration has two trigger modes:
32
32
33
33
AI SRE's IM interaction covers four major platforms. Each platform supports inbound @mentions (webhook), historical message reads, and outbound replies:
34
34
35
-
| Platform | Group @mention| Direct message (DM) | War room auto-diagnosis |
36
-
| --- | --- | --- | --- |
37
-
| Slack | ✅ | ✅ | ✅ |
38
-
| Lark / Feishu | ✅ | ✅ | ✅ |
39
-
| DingTalk | ✅ | ✅ | ✅ |
40
-
| WeCom | ✅ | ✅ | ✅ |
35
+
| Platform | Group @mention| Direct message (DM) | Processing receipt | War room auto-diagnosis |
36
+
| --- | --- | --- | --- | --- |
37
+
| Slack | ✅ | ✅ | ✅ Emoji (plus a native "is working on it" status) | ✅ |
38
+
| Lark / Feishu | ✅ | ✅ | ✅ Emoji | ✅ |
39
+
| DingTalk | ✅ | ✅ | ✅ Built-in emoji | ✅ |
40
+
| WeCom | ✅ | ✅ | ✅ Placeholder message | ✅ |
41
+
42
+
A "processing receipt" is the immediate signal the bot gives once it has your message; how it looks on each platform is described under [Processing Receipt](#processing-receipt).
41
43
42
44
<Note>
43
45
IM interaction requires that you have already connected the corresponding platform's bot in Flashduty (the same IM bot used for alert notifications and collaboration). Complete the bot configuration in Flashduty's IM integration first — only then can AI SRE send and receive messages on that platform.
@@ -65,7 +67,7 @@ In a connected IM group where **Allow AI SRE conversations in IM** is enabled, *
65
67
66
68
<Steps>
67
69
<Steptitle="Detect the mention">
68
-
The platform distinguishes between **group @mentions** and **direct messages**. Once mentioned, AI SRE deduplicates the event and determines whether the message contains an imperative instruction.
70
+
The platform distinguishes between **group @mentions** and **direct messages**. Once mentioned, AI SRE deduplicates the event, puts a "processing" receipt on that message (see [Processing Receipt](#processing-receipt)), and determines whether the message contains an imperative instruction.
69
71
</Step>
70
72
<Steptitle="Bind to a session">
71
73
AI SRE locates a session using "account + platform + chat" as the key: multiple @mentions within the same IM chat are routed to the **same agent session**, preserving context. For a new chat, the platform supplies the opening message of that IM thread and any associated incident as initial context.
@@ -79,6 +81,22 @@ In a connected IM group where **Allow AI SRE conversations in IM** is enabled, *
79
81
The **reply mode** is configurable (off / first / all), controlling whether AI SRE @mentions the person who asked, and whether it replies **in-thread** or in the main channel. In a busy large group, in-thread replies keep the investigation discussion focused without flooding the channel.
80
82
</Tip>
81
83
84
+
### Processing Receipt
85
+
86
+
Once the bot has your message, it first puts an emoji reaction on **the message you sent**, as an immediate "received, working on it" signal. The emoji is picked **at random** from that platform's own pool rather than being fixed:
| Lark / Feishu | Typing, Thinking, Get, CheckMark, OnIt, Hundred, and more |
92
+
| DingTalk | Platform built-in emotions: 思考, 收到, Get, 赞, 比心, OK, and more |
93
+
94
+
On Slack this sits alongside a native "is working on it" status, which the bot also shows.
95
+
96
+
The reaction is removed once the answer is delivered. A single failed removal (a platform hiccup or rate limit) does not leave the emoji on the message forever: the next successful delivery retries the removal with **the exact emoji that was added**, so stuck receipts are eventually cleared.
97
+
98
+
WeCom has no emoji API, so its "processing" receipt is a placeholder message (a random canned acknowledgement such as "On it…"), which the task checklist card and the final answer then replace in place.
99
+
82
100
### How a Reply Reaches the Chat
83
101
84
102
AI SRE's replies in IM are delivered **only through the `reply` tool** — ordinary assistant text is private working text and is never posted automatically. Once you hand it an investigation, every message that appears in the group is one it deliberately submitted, and it is one of two kinds:
| 4 | Observability stack |`observability.md`+ registers the relevant MCP in `tools.md`|
87
+
| 4 | Observability stack |`observability.md`(datasource gaps go through the Monitors-first two-step route, see the Note below; gaps resolved via a direct MCP still register in `tools.md`)|
88
88
| 5 | Runbooks |`runbooks/<topic>.md`, one file per failure class |
In Phase 4, every datasource gap goes through the same **Monitors-first** two-step route: when the scan digest already shows Monitors usage (a datasource or alert rules), the agent names the gap's `type_ident` and hands over the console deep link `/monit/datasource?create=<type_ident>` — opening it lands in Monitors' add-datasource form with that type preselected; `type_ident` values are prometheus / loki / victorialogs / tencent_cls / sls / mysql / postgres / oracle / clickhouse / elasticsearch. Once the datasource is enabled, the built-in `monit-query` capability picks it up automatically — nothing to install. Only when there is no Monitors evidence at all does it fall back to registering the tools you already run as a direct MCP connection (see [MCP](/en/ai-sre/mcp)). `/init` itself never creates or edits a Monitors datasource.
94
+
</Note>
95
+
92
96
<Note>
93
97
In Phase 7, **native kubectl is the primary way to access a cluster**, with k8s-MCP as a fallback (never configure both for the same cluster). This path is BYOC-only: place a **read-only** kubeconfig for each cluster on the Runner (`<workspace>/.kube/<name>.config`); `clusters.md` records only the **path**, never the token.
Copy file name to clipboardExpand all lines: en/ai-sre/insight.mdx
+2Lines changed: 2 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -118,6 +118,8 @@ About **2–3** forward-looking, grounded suggestions. These are **strategic** (
118
118
119
119
When the report surfaces the `expensive-automation` emergent pattern — an automation-entry session that pulls raw listings into context and counts / groups / ranks them itself instead of scripting the aggregation — the next steps include a **Script-first automation rewrite**: ask the agent in chat to rewrite that automation's task prompt script-first, turning the deterministic collection and aggregation into an embedded, tested script.
120
120
121
+
When sessions expose a root-cause gap on the datasource side, the next steps gain another shape: a recommendation to add the matching datasource in Monitors (10 `type_ident` values: prometheus / loki / victorialogs / sls / tencent_cls / elasticsearch / mysql / postgres / oracle / clickhouse). These suggestions carry an evidence bar — the same `type_ident` gap must appear in ≥ 2 distinct sessions, or a single session must give explicit evidence. Once the datasource is added, the built-in `monit-query` capability becomes available automatically, with no skill to install; the report only suggests, never creates or enables a datasource on your behalf, and no longer suggests "install a skill from the Marketplace."
122
+
121
123
<Note>
122
124
Every friction and win must be grounded in at least one real session and carry a verbatim quote as evidence — the report never fabricates sessions, facts, or runbook gaps. If no rankable friction is found, the frictions part displays an empty-state message while the rest of the report still renders — in that case, "the overview itself is the report."
Copy file name to clipboardExpand all lines: en/ai-sre/overview.mdx
+11-2Lines changed: 11 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -59,7 +59,7 @@ AI SRE is available to accounts holding a **valid (enabled)** On-call Pro or hig
59
59
<Accordiontitle="Billing"icon="gift">
60
60
AI SRE is billed on actual usage starting September 16, 2026, 08:00 Beijing time, settled in credits (1 credit = ¥1); usage before that date isn't charged. Metered items are model usage, sandbox online time, and web search. Activation is free, eligible accounts get gift credits each billing period, and you can also buy credit packages (1,000 / 5,000 / 20,000 credits for ¥950 / ¥4,500 / ¥17,000, valid 365 days from payment). Pay-as-you-go is off by default; an admin can turn it on and optionally set a per-period cap. Full terms: [AI SRE Credit Package Purchase Agreement](/en/compliance/ai-sre-credit-package-purchase-agreement).
61
61
62
-
When the gift credits, credit packages, and wallet balance are all used up, AI SRE refuses new conversation requests: the composer is disabled and a **non-dismissible** notice stays pinned above it (the text comes from the server and covers all four shapes — "not activated yet", "turned off", "this period's allowance is spent", and "the wallet is empty"), with a **Go to AI SRE billing** button that opens the billing center (`/wallet/plan?product=ai-sre`). An admin turns on pay-as-you-go, buys a credit package, or tops up the wallet there; once the state recovers, you can continue the conversation.
62
+
When the gift credits, credit packages, and wallet balance are all used up, AI SRE refuses new conversation requests: the composer is disabled and a **non-dismissible** notice stays pinned above it (the text comes from the server and covers all four shapes — "not activated yet", "turned off", "this period's allowance is spent", and "the wallet is empty"), with a **Go to AI SRE billing** button that opens the billing center (`/wallet/plan?product=ai-sre`). An admin turns on pay-as-you-go, buys a credit package, or tops up the wallet there; once the state recovers, you can continue the conversation. A lapsed subscription, or a plan that doesn't include AI SRE, is a different kind of block whose banner button becomes **Go to On-Call subscription** — see [When the subscription lapses](#when-the-subscription-lapses).
Production changes, restarts, rollbacks, and external notifications all require your confirmation before they execute.
@@ -105,7 +105,16 @@ Every request re-checks the account's On-call subscription, so a subscription th
105
105
| The subscription's tier is Pro or higher, but the subscription isn't enabled (expired, or disabled for any other reason) | "This account's On-call subscription is no longer active (expired or disabled), so AI SRE is paused. It resumes automatically once the On-call subscription is renewed." | Renew the On-call subscription |
106
106
| The account has no subscription, or its subscription is below Pro (whether or not it is enabled) | "AI SRE requires the On-call Professional plan. Upgrade the subscription to continue." | Upgrade to Pro or higher |
107
107
108
-
The notice also appears in the console chat page (pinned above the composer, with a **Go to AI SRE billing** entry on its right), in IM sessions, in automation run records, and in A2A call errors; the credits overview reports the same sentence as `blocked_reason=no_license`. Recovery needs no re-activation: once the renewal or upgrade lands, the next request is admitted automatically.
108
+
The notice also appears in the console chat page, in IM sessions, in automation run records, and in A2A call errors; the credits overview reports the same sentence as `blocked_reason=no_license`.
109
+
110
+
On the console chat page it stays pinned above the composer, and the button on its right **follows the blocking reason**:
111
+
112
+
| Blocking reason | Button | Destination |
113
+
|-----------------|--------|-------------|
114
+
| The subscription lapsed, or the plan doesn't include AI SRE (`blocked_reason=no_license` — the two cases on this page) |**Go to On-Call subscription**| The page that sells the On-call subscription, `/wallet/plan?product=oncall`|
115
+
| Everything else (not activated yet, turned off, this period's allowance spent, wallet empty) |**Go to AI SRE billing**| The page that sells AI SRE credits, `/wallet/plan?product=ai-sre`|
116
+
117
+
A subscription block never points at AI SRE billing — that page sells AI SRE credits, which fixes nothing about the subscription. Recovery needs no re-activation: once the renewal or upgrade lands, the next request is admitted automatically.
<Updatelabel="2026-09-23"description="🤖 AI SRE queries Monitors data sources directly, similar-incident merge analysis, configurable Jira status mapping, and more">
8
+
9
+
### AI SRE queries Monitors data sources directly
10
+
11
+
As long as the account has an enabled Monitors data source, AI SRE can query it directly through the alerting engine (10 types, e.g. prometheus, loki, mysql)—no address, no token, no Runner:
12
+
13
+
- Plugin Overview gains a "**Monitors data sources**" chip with a list dialog
14
+
- The MCP connect dialog gains an "**access method**" choice: Monitors data source (in-platform) vs. direct MCP
15
+
- The built-in **monit-query** Skill loads automatically based on your data sources; the former Marketplace template has been delisted
16
+
17
+
See [MCP (External Tools)](/en/ai-sre/mcp) and [Skills](/en/ai-sre/skills).
18
+
19
+
### AI SRE automations: similar-incident merge analysis for on-call incident triggers
20
+
21
+
The on-call incident trigger of Automations gains a "**similar-incident merge analysis**" toggle (on by default): incidents with similarity **≥0.9** (the same standard as intelligent grouping) merge into the in-progress analysis session instead of opening a new one.
22
+
23
+
- Merge window: **10 minutes**, at most **20 incidents** per merge
24
+
- Turning the toggle off restores one session per incident
25
+
26
+
See [Automations](/en/ai-sre/automations).
27
+
28
+
### Artifacts library supports per-person pinning
29
+
30
+
The artifacts library gains per-person pinning: each member can **pin / unpin** frequently used artifacts, and pinned items sort to the top. See [Artifacts](/en/ai-sre/artifacts).
31
+
32
+
### Subagents support cross-environment dispatch
33
+
34
+
A Subagent is no longer confined to the parent session's environment: it can be dispatched to the **cloud sandbox** or **another online BYOC Runner**, using the target environment's tools and credentials. See [Subagents](/en/ai-sre/agents) and [Environments](/en/ai-sre/environments).
35
+
36
+
### AI SRE IM experience improvements
37
+
38
+
-**DingTalk ack emoji receipt**: after a member acknowledges (acks) an incident in IM, AI SRE marks the ack message with an emoji reaction as a receipt; the emoji is picked randomly from each platform's emoji pool and **removed once delivered**
39
+
-**Progress notes as standalone messages**: in-turn progress notes are delivered as separate messages instead of being appended to the same reply
40
+
-**Replies delivered immediately**: AI SRE replies are sent right away, with no duplicate final answer appended when the turn ends
41
+
42
+
See [IM Integration](/en/ai-sre/im).
43
+
44
+
### Monitors data source team-authorization fields land in the console form
45
+
46
+
The data source team-authorization fields (**managing team**`manage_team_id` / **queryable teams**`readonly_team_ids`) are now available in the console form—no longer API-only (the capability first shipped via API on September 16, 2026); members without manage permission get **read-only** access with credential fields masked. See [Data Source Management](/en/monitors/data-sources/data-sources).
47
+
48
+
### Defined behavior for expired On-call subscriptions
49
+
50
+
When the On-call subscription expires:
51
+
52
+
- Alert ingestion **stops**—no new incidents are created
53
+
- Creating escalation spaces, schedules, incidents, war rooms, integrations, and member invitations is **frozen**
54
+
- A **banner notice** shows in the console
55
+
- Service **resumes automatically on renewal**, with no manual steps
56
+
57
+
See [Pricing](/en/platform/pricing).
58
+
59
+
### Jira integration: configurable status mapping; field mapping gains incident attributes
60
+
61
+
-**Status mapping** is now configurable: map the Triggered / Processing / Closed states to Jira statuses individually, or leave blank to use the default mapping
See [Jira Two-Way Sync](/en/on-call/integration/webhooks/jira-sync).
65
+
66
+
### Work items can be assigned to AI SRE
67
+
68
+
The assignees array's `type` now supports `person` or `ai_sre`, so a work item can be assigned to AI SRE; the CLI gains an `--assignee-type` filter. See [CLI](/en/developer/cli).
69
+
70
+
### New alert sources: Checkly and Sumo Logic
71
+
72
+
Two new alert sources join the integration catalog: **Checkly** (monitoring-as-code synthetic monitoring) and **Sumo Logic** (cloud log management and observability). Setup guides: [Checkly](/en/on-call/integration/alert-integration/alert-sources/checkly) and [Sumo Logic](/en/on-call/integration/alert-integration/alert-sources/sumo-logic).
8
73
9
74
### Checkly alert integration added
10
75
@@ -16,7 +81,6 @@ A new Checkly synthetic-monitoring alert integration is available, so Checkly ch
See [Checkly alert integration](/en/on-call/integration/alert-integration/alert-sources/checkly) for the full setup and troubleshooting steps.
19
-
20
84
</Update>
21
85
22
86
<Updatelabel="2026-09-16"description="💰 AI SRE billing goes live, the plugin catalog consolidates onto Overview">
@@ -58,7 +122,7 @@ Data sources gained instance-level authorization: `manage_team_id` (the single m
58
122
- Only three configurations are legal: neither field set (the whole tenant can query and manage); only the managing team set (everyone can query, only the responsible team can manage); both set (query and management are both restricted). Setting `readonly_team_ids` without `manage_team_id` is rejected — a data source with restricted query access must have an explicit responsible team; `manage` always implies `readonly`
59
123
- The data source's creator, the owner account, and the Admin role always keep manage; the list returns only the instances you may query (each row carries `my_perm`, either `manage` or `readonly`), and saving an alert rule checks readonly against every data source the rule references (including the name patterns' matches at save time) — a pure disable is the sole exemption
60
124
- An instance you may not read is dropped from the list, the query entries, and the rule data source picker entirely; an unauthorized operation is refused explicitly without exposing data source details, and changes to the authorization fields are audited
61
-
- The console's data source form does not offer inputs for these two fields yet; an admin sets them through the API
125
+
- The console's data source form does not offer inputs for these two fields yet; an admin sets them through the API**(console form support for these two fields landed on September 23, 2026 — see the September 23 entry)**
62
126
63
127
See [Data Source Management](/en/monitors/data-sources/data-sources), section "Data source permissions (team authorization)".
Copy file name to clipboardExpand all lines: en/monitors/alert-rules/query-result-fields.mdx
+2Lines changed: 2 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -155,6 +155,8 @@ A data source can impose a stricter limit. For example, raw log queries for SLS,
155
155
156
156
When a query exceeds its effective limit, it fails with `too many rows`. Monitors does not return a partial result or evaluate alerts with truncated data.
157
157
158
+
The query preview table on the rule editor page paginates in the frontend: 100 rows per page by default, with options of 100 / 200 / 500, and the paginator is hidden when the results fit in a single page. This pagination affects display only and is independent of the Edge `--alerter.previewMaxRows` limit — a query that exceeds the Edge limit still fails with `too many rows`.
159
+
158
160
<Warning>
159
161
Increasing the Edge row limit does not disable safeguards for response bytes, field size, or total result values. High-cardinality Prometheus queries and large SQL results can still fail on another resource limit.
0 commit comments