Skip to content

Commit 6b1a7c2

Browse files
authored
Merge pull request #964 from flashcatcloud/feat/b23-alert-test
docs: sync batch-23 integration pages to test
2 parents fb3f396 + 3934161 commit 6b1a7c2

16 files changed

Lines changed: 1732 additions & 2 deletions

File tree

‎docs.json‎

Lines changed: 14 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1709,6 +1709,8 @@
17091709
"zh/on-call/integration/alert-integration/alert-sources/duplicati",
17101710
"zh/on-call/integration/alert-integration/alert-sources/robotalp",
17111711
"zh/on-call/integration/alert-integration/alert-sources/infisical",
1712+
"zh/on-call/integration/alert-integration/alert-sources/powerjob",
1713+
"zh/on-call/integration/alert-integration/alert-sources/dolphinscheduler",
17121714
"zh/on-call/integration/alert-integration/alert-sources/fivetran",
17131715
"zh/on-call/integration/alert-integration/alert-sources/coralogix",
17141716
"zh/on-call/integration/alert-integration/alert-sources/uptimeobserver",
@@ -1995,7 +1997,11 @@
19951997
"zh/on-call/integration/change-integration/expo-eas",
19961998
"zh/on-call/integration/change-integration/env0",
19971999
"zh/on-call/integration/change-integration/netbox",
1998-
"zh/on-call/integration/change-integration/nautobot"
2000+
"zh/on-call/integration/change-integration/nautobot",
2001+
"zh/on-call/integration/change-integration/gitee",
2002+
"zh/on-call/integration/change-integration/zadig",
2003+
"zh/on-call/integration/change-integration/yunxiao-appstack",
2004+
"zh/on-call/integration/change-integration/apollo"
19992005
]
20002006
},
20012007
{
@@ -3329,6 +3335,8 @@
33293335
"en/on-call/integration/alert-integration/alert-sources/duplicati",
33303336
"en/on-call/integration/alert-integration/alert-sources/robotalp",
33313337
"en/on-call/integration/alert-integration/alert-sources/infisical",
3338+
"en/on-call/integration/alert-integration/alert-sources/powerjob",
3339+
"en/on-call/integration/alert-integration/alert-sources/dolphinscheduler",
33323340
"en/on-call/integration/alert-integration/alert-sources/fivetran",
33333341
"en/on-call/integration/alert-integration/alert-sources/coralogix",
33343342
"en/on-call/integration/alert-integration/alert-sources/uptimeobserver",
@@ -3615,7 +3623,11 @@
36153623
"en/on-call/integration/change-integration/expo-eas",
36163624
"en/on-call/integration/change-integration/env0",
36173625
"en/on-call/integration/change-integration/netbox",
3618-
"en/on-call/integration/change-integration/nautobot"
3626+
"en/on-call/integration/change-integration/nautobot",
3627+
"en/on-call/integration/change-integration/gitee",
3628+
"en/on-call/integration/change-integration/zadig",
3629+
"en/on-call/integration/change-integration/yunxiao-appstack",
3630+
"en/on-call/integration/change-integration/apollo"
36193631
]
36203632
},
36213633
{
Lines changed: 157 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,157 @@
1+
---
2+
title: "Apache DolphinScheduler Alert Integration"
3+
description: "Send workflow instance failure, success, and timeout notifications from DolphinScheduler's HTTP alert plugin to Flashduty On-call."
4+
keywords: ["alert integration", "DolphinScheduler", "workflow", "Webhook", "job scheduling"]
5+
---
6+
7+
When a workflow instance ends or times out, Apache DolphinScheduler sends a notification to the alert instances in the selected alarm group. If you point an alert instance of the **HTTP** plugin at the Flashduty push URL, a failed workflow instance creates one Critical alert in Flashduty, and the alert recovers when the same instance is re-run and succeeds.
8+
9+
This integration was verified against DolphinScheduler 3.4.3. You must fill in **Body** as shown in [Configure the body](#configure-the-body), or DolphinScheduler sends no alert content.
10+
11+
<div className="hide">
12+
13+
## In Flashduty On-call
14+
---
15+
16+
You can get the integration push URL in either of two ways. Pick one.
17+
18+
### Use a dedicated integration
19+
20+
1. Go to the Flashduty console, select **Channels**, and open a channel
21+
2. Select **Settings** → **Integration data** → **Dedicated integrations**, then click **Add an integration**
22+
3. Select **Dolphinscheduler**, then click **Save**
23+
4. Open the new integration card and copy the **push URL**
24+
25+
### Use a shared integration
26+
27+
1. Go to the Flashduty console and select **Integration Center → Alert events**
28+
2. Select **Dolphinscheduler** and enter an integration name
29+
3. Configure the default route and pick a channel; you can add more rules under **Routes** after it is created
30+
4. Click **Save** and copy the generated **push URL**
31+
32+
</div>
33+
34+
## In DolphinScheduler
35+
---
36+
37+
<Steps>
38+
<Step title="Create an HTTP alarm instance">
39+
40+
1. Sign in to DolphinScheduler with an account that has permission, go to **Security → Alarm Instance Manage**, and click **Create Alarm Instance**
41+
2. For **Select plugin** choose `Http`, and enter an **Alarm instance name**
42+
3. Fill in the plugin parameters:
43+
44+
| Parameter | Value |
45+
| :--- | :--- |
46+
| URL | The full Flashduty push URL, including `integration_key` |
47+
| Request Type | `POST` |
48+
| Headers | Leave empty |
49+
| Body | `{"content":"${msg}"}` |
50+
| Content Type | `application/json` |
51+
| Timeout(s) | Default 120 |
52+
53+
4. Save. The DolphinScheduler server must be able to reach Flashduty on the public internet.
54+
55+
<a id="configure-the-body"></a>
56+
57+
**Configure the body**: the HTTP plugin sends only what you put in **Body**, and replaces `${msg}` inside a string value of the body with the alert content. Without `${msg}`, Flashduty receives no alert content. The value of `content` must be `"${msg}"`; other keys are ignored. After replacement, `content` is a JSON string that holds an array of alert objects, and Flashduty parses it a second time.
58+
59+
</Step>
60+
61+
<Step title="Create an alarm group">
62+
63+
1. Go to **Security → Alarm Group Manage** and click **Create Alarm Group**
64+
2. Enter an **Alert Group Name**, select the alarm instance from the previous step under **Alarm Plugin Instance**, and save
65+
66+
</Step>
67+
68+
<Step title="Select the alarm group and notification strategy on the workflow">
69+
70+
When you start a workflow, or set a schedule for it:
71+
72+
1. Select the alarm group from the previous step for **Alarm Group**
73+
2. Set **Notification Strategy** to **All** (`ALL`). With **Failure**, only failures are sent and alerts in Flashduty never recover
74+
75+
</Step>
76+
77+
<Step title="Verify">
78+
79+
In **Alarm Instance Manage**, click **Test Send** on the alarm instance and confirm Flashduty shows one Info alert titled `DolphinScheduler test notification`. Then make a task in a workflow fail (for example a Shell task that runs `exit 1`) and confirm a Critical alert titled `DolphinScheduler workflow failed: <workflow instance name>` appears.
80+
81+
</Step>
82+
</Steps>
83+
84+
## Recovery and auto-close
85+
---
86+
87+
DolphinScheduler sends a notification according to the state the workflow instance ends in:
88+
89+
| Workflow instance state | What Flashduty does |
90+
| :--- | :--- |
91+
| `FAILURE` | Triggers a Critical alert |
92+
| `SUCCESS` | Recovers the alert of the same instance |
93+
| `STOP`, `PAUSE`, and other states | Ignored. A manual stop or pause is neither a failure nor a recovery |
94+
95+
When you re-run a failed instance with **Recovery Failed** (`START_FAILURE_TASK_PROCESS`), the instance ID stays the same and `runTimes` increases by 1. If the re-run fails again, the same alert is updated. If it succeeds and the notification strategy is **All**, the alert recovers.
96+
97+
Workflow and task timeout alerts are sent once and never recover. If nobody re-runs a failed workflow, no recovery is sent either. In the channel that receives this integration, turn on [auto-close](/en/on-call/channel/create-edit). We suggest 12 hours; adjust to how quickly your team handles failed workflows.
98+
99+
Sub-workflow states are not notified on their own.
100+
101+
## Alert types
102+
---
103+
104+
The `content` array of one request can hold several alert objects. Each object becomes one Flashduty alert.
105+
106+
- **Workflow instance failure or success**: the title is `DolphinScheduler workflow failed: <workflow instance name>` or `DolphinScheduler workflow succeeded: <workflow instance name>`
107+
- **Workflow timeout**: the title is `DolphinScheduler workflow timeout: <workflow instance name>`, Warning severity
108+
- **Task timeout**: the title is `DolphinScheduler task timeout: <task name>`, Warning severity
109+
- **Test send**: the test message of an alarm instance is fixed. Flashduty recognizes it and creates a separate Info alert; every press is a new alert and never merges with or closes a real alert. Close it by hand
110+
111+
If the alarm group also receives alerts that carry no workflow information (for example a service-down alert), Flashduty rejects them with an invalid-parameter error. To keep failed sends out of DolphinScheduler, create a separate alarm group for this integration.
112+
113+
## Alert Key
114+
---
115+
116+
| Alert type | Alert Key |
117+
| :--- | :--- |
118+
| Workflow instance failure or success | project code `projectCode` + workflow instance ID `workflowInstanceId` |
119+
| Workflow timeout | project code + workflow instance ID, plus the fixed marker `timeout` |
120+
| Task timeout | project code + workflow instance ID + task code `taskCode`, plus the fixed marker `timeout` |
121+
122+
Changes to the workflow instance name, state, run count, or times do not change the Alert Key. A failure and the later success of the same instance share one Alert Key, so the success recovers the failure. An alert object without `projectCode` or `workflowInstanceId` makes the whole request be rejected.
123+
124+
## Status and severity
125+
---
126+
127+
| Source | Status | Severity |
128+
| :--- | :--- | :--- |
129+
| `FAILURE` | Triggered | Critical |
130+
| `SUCCESS` | Recovered | Keeps the severity it was triggered with |
131+
| Timeout | Triggered | Warning |
132+
133+
## Labels
134+
---
135+
136+
| Label | Source |
137+
| :--- | :--- |
138+
| `source` | Always `dolphinscheduler` |
139+
| `check` | Workflow instance name (`workflow instance <ID>` when missing) |
140+
| `resource` | Project name `projectName` |
141+
| `project_code` / `project_name` | Project code and name |
142+
| `workflow_instance_id` / `workflow_instance_name` | Workflow instance ID and name |
143+
| `workflow_definition_code` | Workflow definition code |
144+
| `command_type` | How it was started, for example `START_PROCESS`, `START_FAILURE_TASK_PROCESS` |
145+
| `run_times` | Run count |
146+
| `workflow_status` | Workflow instance state |
147+
| `workflow_host` | Master address that ran the workflow |
148+
| `event` / `warn_level` | Event and level of a timeout alert |
149+
| `task_code` / `task_name` | Task code and name (task timeout) |
150+
151+
## Troubleshooting
152+
---
153+
154+
- **Flashduty receives no events**: confirm the workflow was started with an alarm group and a notification strategy other than **None**, that the alarm group contains the alarm instance, and that the push URL is complete and includes `integration_key`. **Test Send** in **Alarm Instance Manage** checks the network path
155+
- **Flashduty returns an invalid-parameter error**: usually **Body** is not `{"content":"${msg}"}`, or an alert without workflow information (for example service-down) was received
156+
- **An alert never closes**: the failed instance was not re-run successfully, or the notification strategy is not **All**. Turn on auto-close for the channel
157+
- **No success notifications**: the notification strategy is **Failure**
Lines changed: 122 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,122 @@
1+
---
2+
title: "PowerJob alert integration"
3+
description: "Send failed job instances and failed workflow instances from PowerJob to Flashduty On-call through a user WebHook alarm."
4+
keywords: ["alert integration", "PowerJob", "scheduled jobs", "workflow", "webhook", "job scheduling"]
5+
---
6+
7+
When a job instance or a workflow instance fails, PowerJob posts a JSON alarm to the WebHook address of each notified user selected on the job (or workflow). Set that address to the Flashduty push URL, and every failed run creates one Flashduty alert with Critical severity.
8+
9+
PowerJob alarms only on failure and sends nothing when a run succeeds, so the alert does not recover on its own. Turn on auto-close in the channel, see [Alerts do not recover](#alerts-do-not-recover).
10+
11+
<div className="hide">
12+
13+
## In Flashduty On-call
14+
---
15+
16+
You can get the push URL in either of the following ways.
17+
18+
### Use a dedicated integration
19+
20+
1. In the Flashduty console, select **Channels** and open a channel
21+
2. Select **Settings** → **Integration data** → **Dedicated integration**, then click **Add an integration**
22+
3. Select **PowerJob** and click **Save**
23+
4. Open the generated integration card and copy the **push URL**
24+
25+
### Use a shared integration
26+
27+
1. In the Flashduty console, select **Integration Center → Alert events**
28+
2. Select **PowerJob** and enter an integration name
29+
3. Configure the default route and select a channel; you can add more rules under **Routes** after the integration is created
30+
4. Click **Save** and copy the generated **push URL**
31+
32+
</div>
33+
34+
## In PowerJob
35+
---
36+
37+
PowerJob sends alarms to users: first set a WebHook address on a user, then select that user on the job or workflow. This integration is verified against PowerJob 5.1.7.
38+
39+
<Steps>
40+
<Step title="Set the WebHook address on a user">
41+
42+
1. Sign in to the PowerJob console and open the **Personal Info** page
43+
2. Enter the complete Flashduty push URL (including `integration_key`) in the **WebHook** field and save
44+
45+
If the address has no `http://` or `https://` prefix, PowerJob adds `http://`, so enter the full address starting with `https://`. The PowerJob server must be able to reach Flashduty on the public internet.
46+
47+
</Step>
48+
49+
<Step title="Select the notified users on a job or workflow">
50+
51+
1. Edit the job you want to monitor. In **Alarm config**, select the user from the previous step in **Alarm receiver(s)**, then save
52+
2. Do the same for each workflow you want to monitor
53+
54+
Only selected users that have a WebHook receive the alarm. PowerJob sends one request for each such user.
55+
56+
</Step>
57+
58+
<Step title="Verify">
59+
60+
PowerJob has no test button. Make a job fail, for example with the built-in `StandaloneProcessorDemo` processor and the job parameter set to `failed`. Run it and confirm Flashduty shows one Critical alert titled `PowerJob job failed: <job name>`.
61+
62+
</Step>
63+
</Steps>
64+
65+
## Alerts do not recover
66+
---
67+
68+
PowerJob sends no recovery notification, and it does not alarm when a job later succeeds. In the channel that receives this integration, turn on [auto-close](/en/on-call/channel/create-edit). We suggest 12 hours; adjust to how quickly your team handles failed jobs.
69+
70+
Each failed run is a separate alert. A job that is scheduled often (for example every minute) and keeps failing creates many alerts; use Flashduty alert grouping to reduce the noise.
71+
72+
## Job failures and workflow failures
73+
---
74+
75+
The integration tells the two alarms apart by their content:
76+
77+
- **Job failure**: contains `jobId` and `instanceId`. The title is `PowerJob job failed: <job name>`, or `job <jobId>` when the name is missing
78+
- **Workflow failure**: contains `workflowId` and `wfInstanceId`. The workflow alarm of PowerJob 5.1.7 carries no workflow name, so the title is `PowerJob workflow failed: workflow <workflowId>`; find the workflow in the PowerJob console by that ID
79+
80+
If a job inside a workflow fails and that job also has notified users, PowerJob sends both a job failure alarm and a workflow failure alarm. They are two separate alerts in Flashduty.
81+
82+
The alert description contains the execution `result` reported by PowerJob (cut at 1000 bytes). Job parameters, instance parameters, and the processor info (for example the body of a shell script) may hold secrets, so Flashduty does not read those fields.
83+
84+
## Alert Key
85+
---
86+
87+
| Alarm type | Alert Key |
88+
| :--- | :--- |
89+
| Job failure | app ID `appId` + job ID `jobId` + job instance ID `instanceId` |
90+
| Workflow failure | workflow ID `workflowId` + workflow instance ID `wfInstanceId` |
91+
92+
Every run has a new instance ID, so each failure is a new alert and the same run never alerts twice. Changes to the job name, result, or times do not change the Alert Key. A request without these fields is rejected.
93+
94+
Instance IDs are long integers above 2^53; Flashduty keeps every digit of the original text.
95+
96+
## Status and severity
97+
---
98+
99+
PowerJob alarms carry no severity field and are sent only on failure, so every alert is in the triggered state with Critical severity.
100+
101+
## Labels
102+
---
103+
104+
| Label | Source |
105+
| :--- | :--- |
106+
| `source` | Always `powerjob` |
107+
| `kind` | `job` or `workflow` |
108+
| `check` | Job name (`job <jobId>` when missing), or `workflow <workflowId>` |
109+
| `resource` | `app <appId>` |
110+
| `app_id` | App ID |
111+
| `job_id` / `job_name` / `instance_id` | Job ID, job name, and job instance ID (job failure) |
112+
| `workflow_id` / `wf_instance_id` | Workflow ID and workflow instance ID (workflow failure) |
113+
| `time_expression` | Time expression, for example a CRON expression |
114+
| `task_tracker_address` | Address of the TaskTracker that ran the job (job failure) |
115+
116+
## Troubleshooting
117+
---
118+
119+
- **Flashduty receives no event**: confirm the failed job or workflow has notified users selected and that user has a WebHook; confirm the PowerJob server can reach Flashduty and that the push URL is complete with `integration_key`. When delivery fails, PowerJob only logs `[WebHookAlarmService] invoke webhook ... failed` on the server and does not retry
120+
- **Flashduty returns an invalid-parameter error**: the request lacks the fields listed under Alert Key, or it is not a PowerJob JSON alarm
121+
- **The alert never closes**: this is expected; turn on auto-close in the channel
122+
- **No workflow name**: see above, the workflow alarm of PowerJob 5.1.7 carries no name

‎en/on-call/integration/alert-integration/alert-sources/prometheus.mdx‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -150,6 +150,14 @@ VMP's general webhook sends alerts in the Alertmanager format, so it uses this i
150150

151151
</div>
152152

153+
<Note>
154+
Volcengine is moving VMP alert notification to the Cloud Monitor alert center. According to Volcengine's announcement, after October 12, 2026 alert configurations that have not been migrated are migrated automatically, and the old VMP alert notification APIs are taken offline gradually. After the migration, VMP alerts can reach Flashduty through the existing [Volcengine Cloud Monitor Alert Events](/en/on-call/integration/alert-integration/alert-sources/volcengine-cm-metrics) integration. Alternatively, use a **Notification Callback** template in the alert center that renders an Alertmanager-format webhook body, and push it to this integration.
155+
</Note>
156+
157+
<Tip>
158+
VMP's `alertname` label is the alert rule's UUID, so the default alert title is a UUID. The rule name is in the `alerting_rule_name` label; set this integration's **Title Rule** to use `$alerting_rule_name`.
159+
</Tip>
160+
153161
## Severity Mapping
154162
---
155163

0 commit comments

Comments
 (0)