From 43bb05a1baee3c713732acea497d153e465cfbbc Mon Sep 17 00:00:00 2001 From: Pau Capdevila Date: Wed, 29 Jul 2026 23:00:30 +0200 Subject: [PATCH 1/2] docs(troubleshooting): add support diagnostics page Signed-off-by: Pau Capdevila --- docs/troubleshooting/.pages | 1 + docs/troubleshooting/support_diagnostics.md | 47 +++++++++++++++++++++ 2 files changed, 48 insertions(+) create mode 100644 docs/troubleshooting/support_diagnostics.md diff --git a/docs/troubleshooting/.pages b/docs/troubleshooting/.pages index 4a3c74c2..df2ae4ba 100644 --- a/docs/troubleshooting/.pages +++ b/docs/troubleshooting/.pages @@ -4,3 +4,4 @@ nav: - LLDP Neighbors: lldp.md - Switch Agent: agent.md - ... + - Support Diagnostics: support_diagnostics.md diff --git a/docs/troubleshooting/support_diagnostics.md b/docs/troubleshooting/support_diagnostics.md new file mode 100644 index 00000000..e2483575 --- /dev/null +++ b/docs/troubleshooting/support_diagnostics.md @@ -0,0 +1,47 @@ +# Support Diagnostics + +If the steps in the rest of this section don't resolve the issue, collect the +following before reaching out to support. The more of this is included up +front, the faster the issue can be localized. + +## What to include + +- Your `hhfab`/`hhfabctl` version and a summary of your topology (gateway + present, external peering, VPCs in use). +- A plain description of the symptom and roughly when it started. +- Anything that changed around that time - a config push, upgrade, reboot, + or node/server move - even if you're not sure it's related. +- Which specific switches or nodes are affected, if you've already narrowed + it down. +- Any diagnostics you've already collected yourself. Include the raw output + rather than a summary, so we don't have to re-derive what you've already + found. + +## Collecting a support bundle + +`hhfabctl` is installed on the control node as a `kubectl` plugin, so it's +invoked as `kubectl hhfab`, not as a standalone command: + +```console +core@control-1 ~ $ kubectl hhfab support dump -y +``` + +This produces a single timestamped `.hhs` file containing cluster resources +(Fabricator, Agent, Connection, VPC, and related objects) and pod logs. +Secrets are redacted, but the bundle still reflects your deployment's real +topology and state. + +## Switch-level diagnostics + +For issues that look like a dataplane or hardware problem rather than a +control-plane one (for example, traffic not forwarding despite correct BGP/EVPN +state), support may also ask for a `show techsupport` capture from specific +switches. This is a per-device SONiC command - if asked, run it only on the +switches identified as relevant rather than the whole fabric, both to keep +the capture a manageable size and because a comparison against a switch that +*is* behaving correctly is often the most useful artifact. + +!!! note + Support bundles and switch dumps reflect your real deployment and should + be treated as confidential. Share them only through the channel your + support contact provides, not in a public issue or channel. From 1993ad8ebe64fb52457123a672a449b0ed5463f8 Mon Sep 17 00:00:00 2001 From: Pau Capdevila Date: Sat, 26 Sep 2026 11:45:09 +0200 Subject: [PATCH 2/2] docs: add agent-log guidance, fix raw-output bullet phrasing Adds a section on collecting /var/log/agent.log for BGP/BFD/interface issues that won't stabilize, learned from a real triage where the support bundle's Agent CR status looked converged the whole time despite active flapping (githedgehog/docs#353). Also rewords the raw-output bullet per review: encourage raw output plus the reporter's own summary, not raw output instead of one. Co-Authored-By: Claude Sonnet 5 Signed-off-by: Pau Capdevila --- docs/troubleshooting/support_diagnostics.md | 20 ++++++++++++++++++-- 1 file changed, 18 insertions(+), 2 deletions(-) diff --git a/docs/troubleshooting/support_diagnostics.md b/docs/troubleshooting/support_diagnostics.md index e2483575..fc4e35bf 100644 --- a/docs/troubleshooting/support_diagnostics.md +++ b/docs/troubleshooting/support_diagnostics.md @@ -14,8 +14,8 @@ front, the faster the issue can be localized. - Which specific switches or nodes are affected, if you've already narrowed it down. - Any diagnostics you've already collected yourself. Include the raw output - rather than a summary, so we don't have to re-derive what you've already - found. + so we can reach our own conclusions and rule out other causes, plus your + own summary or theory if you have one. ## Collecting a support bundle @@ -31,6 +31,22 @@ This produces a single timestamped `.hhs` file containing cluster resources Secrets are redacted, but the bundle still reflects your deployment's real topology and state. +## Switch agent logs + +For issues where BGP, BFD, or interface state won't stabilize or keeps +flapping, the support bundle's Agent CR status can look fully converged even +while the problem is ongoing - it only reflects the most recent apply, not +churn happening between applies. `/var/log/agent.log` on the affected +switches is the artifact that actually shows this. It's a plain file, not a +`kubectl` or `sonic-cli` command: + +```console +admin@switch:~$ cat /var/log/agent.log +``` + +Include the whole file, or at least the span covering when the issue was +observed, rather than a filtered excerpt. + ## Switch-level diagnostics For issues that look like a dataplane or hardware problem rather than a