|
| 1 | +--- |
| 2 | +title: "Apache DolphinScheduler Alert Integration" |
| 3 | +description: "Send workflow instance failure, success, and timeout notifications from DolphinScheduler's HTTP alert plugin to Flashduty On-call." |
| 4 | +keywords: ["alert integration", "DolphinScheduler", "workflow", "Webhook", "job scheduling"] |
| 5 | +--- |
| 6 | + |
| 7 | +When a workflow instance ends or times out, Apache DolphinScheduler sends a notification to the alert instances in the selected alarm group. If you point an alert instance of the **HTTP** plugin at the Flashduty push URL, a failed workflow instance creates one Critical alert in Flashduty, and the alert recovers when the same instance is re-run and succeeds. |
| 8 | + |
| 9 | +This integration was verified against DolphinScheduler 3.4.3. You must fill in **Body** as shown in [Configure the body](#configure-the-body), or DolphinScheduler sends no alert content. |
| 10 | + |
| 11 | +<div className="hide"> |
| 12 | + |
| 13 | +## In Flashduty On-call |
| 14 | +--- |
| 15 | + |
| 16 | +You can get the integration push URL in either of two ways. Pick one. |
| 17 | + |
| 18 | +### Use a dedicated integration |
| 19 | + |
| 20 | +1. Go to the Flashduty console, select **Channels**, and open a channel |
| 21 | +2. Select **Settings** → **Integration data** → **Dedicated integrations**, then click **Add an integration** |
| 22 | +3. Select **Dolphinscheduler**, then click **Save** |
| 23 | +4. Open the new integration card and copy the **push URL** |
| 24 | + |
| 25 | +### Use a shared integration |
| 26 | + |
| 27 | +1. Go to the Flashduty console and select **Integration Center → Alert events** |
| 28 | +2. Select **Dolphinscheduler** and enter an integration name |
| 29 | +3. Configure the default route and pick a channel; you can add more rules under **Routes** after it is created |
| 30 | +4. Click **Save** and copy the generated **push URL** |
| 31 | + |
| 32 | +</div> |
| 33 | + |
| 34 | +## In DolphinScheduler |
| 35 | +--- |
| 36 | + |
| 37 | +<Steps> |
| 38 | +<Step title="Create an HTTP alarm instance"> |
| 39 | + |
| 40 | +1. Sign in to DolphinScheduler with an account that has permission, go to **Security → Alarm Instance Manage**, and click **Create Alarm Instance** |
| 41 | +2. For **Select plugin** choose `Http`, and enter an **Alarm instance name** |
| 42 | +3. Fill in the plugin parameters: |
| 43 | + |
| 44 | +| Parameter | Value | |
| 45 | +| :--- | :--- | |
| 46 | +| URL | The full Flashduty push URL, including `integration_key` | |
| 47 | +| Request Type | `POST` | |
| 48 | +| Headers | Leave empty | |
| 49 | +| Body | `{"content":"${msg}"}` | |
| 50 | +| Content Type | `application/json` | |
| 51 | +| Timeout(s) | Default 120 | |
| 52 | + |
| 53 | +4. Save. The DolphinScheduler server must be able to reach Flashduty on the public internet. |
| 54 | + |
| 55 | +<a id="configure-the-body"></a> |
| 56 | + |
| 57 | +**Configure the body**: the HTTP plugin sends only what you put in **Body**, and replaces `${msg}` inside a string value of the body with the alert content. Without `${msg}`, Flashduty receives no alert content. The value of `content` must be `"${msg}"`; other keys are ignored. After replacement, `content` is a JSON string that holds an array of alert objects, and Flashduty parses it a second time. |
| 58 | + |
| 59 | +</Step> |
| 60 | + |
| 61 | +<Step title="Create an alarm group"> |
| 62 | + |
| 63 | +1. Go to **Security → Alarm Group Manage** and click **Create Alarm Group** |
| 64 | +2. Enter an **Alert Group Name**, select the alarm instance from the previous step under **Alarm Plugin Instance**, and save |
| 65 | + |
| 66 | +</Step> |
| 67 | + |
| 68 | +<Step title="Select the alarm group and notification strategy on the workflow"> |
| 69 | + |
| 70 | +When you start a workflow, or set a schedule for it: |
| 71 | + |
| 72 | +1. Select the alarm group from the previous step for **Alarm Group** |
| 73 | +2. Set **Notification Strategy** to **All** (`ALL`). With **Failure**, only failures are sent and alerts in Flashduty never recover |
| 74 | + |
| 75 | +</Step> |
| 76 | + |
| 77 | +<Step title="Verify"> |
| 78 | + |
| 79 | +In **Alarm Instance Manage**, click **Test Send** on the alarm instance and confirm Flashduty shows one Info alert titled `DolphinScheduler test notification`. Then make a task in a workflow fail (for example a Shell task that runs `exit 1`) and confirm a Critical alert titled `DolphinScheduler workflow failed: <workflow instance name>` appears. |
| 80 | + |
| 81 | +</Step> |
| 82 | +</Steps> |
| 83 | + |
| 84 | +## Recovery and auto-close |
| 85 | +--- |
| 86 | + |
| 87 | +DolphinScheduler sends a notification according to the state the workflow instance ends in: |
| 88 | + |
| 89 | +| Workflow instance state | What Flashduty does | |
| 90 | +| :--- | :--- | |
| 91 | +| `FAILURE` | Triggers a Critical alert | |
| 92 | +| `SUCCESS` | Recovers the alert of the same instance | |
| 93 | +| `STOP`, `PAUSE`, and other states | Ignored. A manual stop or pause is neither a failure nor a recovery | |
| 94 | + |
| 95 | +When you re-run a failed instance with **Recovery Failed** (`START_FAILURE_TASK_PROCESS`), the instance ID stays the same and `runTimes` increases by 1. If the re-run fails again, the same alert is updated. If it succeeds and the notification strategy is **All**, the alert recovers. |
| 96 | + |
| 97 | +Workflow and task timeout alerts are sent once and never recover. If nobody re-runs a failed workflow, no recovery is sent either. In the channel that receives this integration, turn on [auto-close](/en/on-call/channel/create-edit). We suggest 12 hours; adjust to how quickly your team handles failed workflows. |
| 98 | + |
| 99 | +Sub-workflow states are not notified on their own. |
| 100 | + |
| 101 | +## Alert types |
| 102 | +--- |
| 103 | + |
| 104 | +The `content` array of one request can hold several alert objects. Each object becomes one Flashduty alert. |
| 105 | + |
| 106 | +- **Workflow instance failure or success**: the title is `DolphinScheduler workflow failed: <workflow instance name>` or `DolphinScheduler workflow succeeded: <workflow instance name>` |
| 107 | +- **Workflow timeout**: the title is `DolphinScheduler workflow timeout: <workflow instance name>`, Warning severity |
| 108 | +- **Task timeout**: the title is `DolphinScheduler task timeout: <task name>`, Warning severity |
| 109 | +- **Test send**: the test message of an alarm instance is fixed. Flashduty recognizes it and creates a separate Info alert; every press is a new alert and never merges with or closes a real alert. Close it by hand |
| 110 | + |
| 111 | +If the alarm group also receives alerts that carry no workflow information (for example a service-down alert), Flashduty rejects them with an invalid-parameter error. To keep failed sends out of DolphinScheduler, create a separate alarm group for this integration. |
| 112 | + |
| 113 | +## Alert Key |
| 114 | +--- |
| 115 | + |
| 116 | +| Alert type | Alert Key | |
| 117 | +| :--- | :--- | |
| 118 | +| Workflow instance failure or success | project code `projectCode` + workflow instance ID `workflowInstanceId` | |
| 119 | +| Workflow timeout | project code + workflow instance ID, plus the fixed marker `timeout` | |
| 120 | +| Task timeout | project code + workflow instance ID + task code `taskCode`, plus the fixed marker `timeout` | |
| 121 | + |
| 122 | +Changes to the workflow instance name, state, run count, or times do not change the Alert Key. A failure and the later success of the same instance share one Alert Key, so the success recovers the failure. An alert object without `projectCode` or `workflowInstanceId` makes the whole request be rejected. |
| 123 | + |
| 124 | +## Status and severity |
| 125 | +--- |
| 126 | + |
| 127 | +| Source | Status | Severity | |
| 128 | +| :--- | :--- | :--- | |
| 129 | +| `FAILURE` | Triggered | Critical | |
| 130 | +| `SUCCESS` | Recovered | Keeps the severity it was triggered with | |
| 131 | +| Timeout | Triggered | Warning | |
| 132 | + |
| 133 | +## Labels |
| 134 | +--- |
| 135 | + |
| 136 | +| Label | Source | |
| 137 | +| :--- | :--- | |
| 138 | +| `source` | Always `dolphinscheduler` | |
| 139 | +| `check` | Workflow instance name (`workflow instance <ID>` when missing) | |
| 140 | +| `resource` | Project name `projectName` | |
| 141 | +| `project_code` / `project_name` | Project code and name | |
| 142 | +| `workflow_instance_id` / `workflow_instance_name` | Workflow instance ID and name | |
| 143 | +| `workflow_definition_code` | Workflow definition code | |
| 144 | +| `command_type` | How it was started, for example `START_PROCESS`, `START_FAILURE_TASK_PROCESS` | |
| 145 | +| `run_times` | Run count | |
| 146 | +| `workflow_status` | Workflow instance state | |
| 147 | +| `workflow_host` | Master address that ran the workflow | |
| 148 | +| `event` / `warn_level` | Event and level of a timeout alert | |
| 149 | +| `task_code` / `task_name` | Task code and name (task timeout) | |
| 150 | + |
| 151 | +## Troubleshooting |
| 152 | +--- |
| 153 | + |
| 154 | +- **Flashduty receives no events**: confirm the workflow was started with an alarm group and a notification strategy other than **None**, that the alarm group contains the alarm instance, and that the push URL is complete and includes `integration_key`. **Test Send** in **Alarm Instance Manage** checks the network path |
| 155 | +- **Flashduty returns an invalid-parameter error**: usually **Body** is not `{"content":"${msg}"}`, or an alert without workflow information (for example service-down) was received |
| 156 | +- **An alert never closes**: the failed instance was not re-run successfully, or the notification strategy is not **All**. Turn on auto-close for the channel |
| 157 | +- **No success notifications**: the notification strategy is **Failure** |
0 commit comments