-
-
Notifications
You must be signed in to change notification settings - Fork 289
feat(docs): site health + AI-citation fixes (A1, A2, A8 …) #897
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
52 commits
Select commit
Hold shift + click to select a range
f34e09b
fix(docs): add image to Article schema (A1)
dhananjay6561 88f8864
fix(docs): differentiate page title from h1 (A2)
dhananjay6561 52568df
fix(docs): add alt text to images missing it (A8)
dhananjay6561 23c9c94
style(docs): prettier formatting on api-testing-functions/variables
dhananjay6561 0301cd3
style(docs): prettier (2.8.8) on DocItem
dhananjay6561 7f927db
style(docs): format DocItem for CI prettier 3.9.6
dhananjay6561 b4d42f0
feat(docs): specialize doc schema to TechArticle/APIReference (Doc2)
dhananjay6561 563815c
fix(docs): unique meta descriptions for 6 pages (A7)
dhananjay6561 98e2a80
style(docs): prettier 3.9.6 on beta-testing.md
dhananjay6561 ae50868
feat(docs): add structured data to bespoke pages (Doc2)
dhananjay6561 46a9ac3
feat(docs): emit FAQPage schema for FAQ docs (Doc2/AI4)
dhananjay6561 03ffef4
feat(docs): consolidate JSON-LD into one @id-linked entity graph (Doc2)
dhananjay6561 429537c
fix(docs): resolve residual SEO audit items (A8, A9)
dhananjay6561 4ffad60
docs(A5): expand thin SCM PR-agent page with capabilities + related l…
dhananjay6561 f7ae0b9
docs(A5): expand Windows/WSL install page with prerequisites + relate…
dhananjay6561 8f505e6
docs(A4): add "Related Terms" cross-links to all 37 glossary pages
dhananjay6561 3104c08
docs(A4): add "Related" cross-links to running-keploy docs
dhananjay6561 a292fef
docs(A4): add "Related" cross-links to quickstart sample apps
dhananjay6561 96e82dc
docs(A4): add "Related" cross-links to keploy-cloud docs
dhananjay6561 6e6b4be
docs(A4): add "Related" cross-links to keploy-explained docs
dhananjay6561 cf53f7f
docs(A4): add "Related" cross-links to ci-cd docs
dhananjay6561 881f832
docs(A4): add "Related" cross-links to server install + SDK docs
dhananjay6561 a66ad50
docs(AI4): add HowTo schema to CI/CD integration guides
dhananjay6561 668e9df
docs(AI4): add HowTo schema to language SDK install guides
dhananjay6561 78440b9
docs(AI4): add HowTo schema to Linux/Windows install guides
dhananjay6561 409fbff
ci(vale): accept technical terms flagged on changed lines
dhananjay6561 2c06a80
feat(docs): add CollectionPage on hubs, LearningResource on quickstarts
dhananjay6561 be97582
feat(docs): add WebPage/BreadcrumbList/ItemList to application-develo…
dhananjay6561 1e9f10f
feat(docs): emit community-channels ItemList on the home page
dhananjay6561 0295310
refactor(docs): centralize breadcrumb JSON-LD in a shared builder
dhananjay6561 b6848d7
fix(docs): add repo-hosted 1200x630 social card
dhananjay6561 2075ee0
fix(docs): sync SEO title and social-image metadata in the doc theme
dhananjay6561 9a49f30
ci(docs): run the schema-graph verifier in the PR build check
dhananjay6561 8dc2fd4
fix(docs): harden the schema-graph guard
dhananjay6561 197a8e4
fix(docs): correct leadership route in schema and permalink
dhananjay6561 25fcc97
fix(docs): keep FAQ answers readable and drop the Related section
dhananjay6561 1c61728
fix(docs): add trailing slash to the SearchAction target
dhananjay6561 9b744c4
fix(docs): add alt text to remaining v4 images
dhananjay6561 385d444
fix(docs): unique meta descriptions for duplicate pages
dhananjay6561 ea351c1
fix(docs): schema-type accuracy in the doc theme
dhananjay6561 641a546
fix(docs): unique descriptions for the three Node.js sample apps
dhananjay6561 3568541
fix(docs): bring two edited descriptions within the 70-160 band
dhananjay6561 fea6bb4
Merge branch 'main' into feat/ai-citation-health
dhananjay6561 e579122
style(docs): format remarkFaqSchema per prettier 3.9.6
dhananjay6561 1533683
fix(docs): normalize en-dashes to em-dash or hyphen for Vale
dhananjay6561 12a77aa
docs: unlink NDJSON in the public API reference
dhananjay6561 bba636e
revert(docs): undo en-dash edits in non-Vale / pre-debt files
dhananjay6561 53097bb
revert(docs): drop en-dash edits in noIndex v2/v3 pages
dhananjay6561 e41303d
fix(docs): null-guard title before .length in DocItem
dhananjay6561 0506372
fix(docs): add prose lead-in so FAQ Q3 enters FAQPage schema
dhananjay6561 df65cfe
Merge branch 'main' into feat/ai-citation-health
dhananjay6561 87e1949
style(docs): prettier --write on the 8 files this PR touches
dhananjay6561 File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,165 @@ | ||
| #!/usr/bin/env node | ||
| /** | ||
| * Verifies the JSON-LD graph in the built site. | ||
| * | ||
| * The @id-based consolidation only pays off if every bare {"@id": "..."} | ||
| * reference resolves to a node actually defined on the same page -- an | ||
| * unresolved reference is worse than the inline duplicate it replaced, | ||
| * because consumers get a dangling pointer instead of an entity. This walks | ||
| * the built HTML and fails on unparseable JSON-LD, dangling references, or a | ||
| * regression to the generic `Article` type. noindex pages (archived versions) | ||
| * are skipped. Note it verifies JSON validity and @id resolution, not that | ||
| * every emitted `url` actually resolves. | ||
| * | ||
| * Usage: node scripts/verify-schema-graph.js [buildDir] | ||
| */ | ||
| const fs = require("fs"); | ||
| const path = require("path"); | ||
|
|
||
| const buildDir = process.argv[2] || "build"; | ||
|
|
||
| // Pages Docusaurus renders with a `noindex` robots meta (archived versions | ||
| // carrying `noIndex: true`, etc.) are skipped: they keep their own legacy | ||
| // schema copies and are never served to crawlers. Detected from the built HTML | ||
| // rather than a hard-coded version list, so it can't drift when a version is | ||
| // archived. | ||
| const NOINDEX_RE = /<meta[^>]+name="robots"[^>]+content="[^"]*noindex/i; | ||
|
|
||
| function findHtml(dir, out = []) { | ||
| for (const entry of fs.readdirSync(dir, {withFileTypes: true})) { | ||
| const full = path.join(dir, entry.name); | ||
| if (entry.isDirectory()) { | ||
| findHtml(full, out); | ||
| } else if (entry.name === "index.html") { | ||
| out.push(full); | ||
| } | ||
| } | ||
| return out; | ||
| } | ||
|
|
||
| // The build inlines JSON-LD into <script type="application/ld+json">. The | ||
| // <Head>/Helmet emitters HTML-escape entities while the raw-text body <script> | ||
| // (remark FAQ plugin) does not, so unescape defensively before parsing -- | ||
| // unescaping already-clean JSON is a no-op here. | ||
| const SCRIPT_RE = | ||
| /<script[^>]*type="application\/ld\+json"[^>]*>([\s\S]*?)<\/script>/g; | ||
|
|
||
| // A page is its own WebPage: the idiomatic `mainEntityOfPage: | ||
| // {"@type":"WebPage","@id":<pageUrl>}` points at the document itself, which has | ||
| // no separate full node. Seed `defined` with the page's own URL so that | ||
| // self-reference resolves, while a typed ref to any *other* undefined @id is | ||
| // still caught as dangling. Use og:url, not the canonical <link>: a few docs | ||
| // set a cross-site canonical (e.g. to the blog), but og:url is always the | ||
| // page's own trailing-slash URL, which is what the schema @id derives from. | ||
| const OG_URL_RE = /<meta[^>]+property="og:url"[^>]+content="([^"]+)"/i; | ||
|
|
||
| function unescapeHtml(s) { | ||
| return s | ||
| .replace(/"/g, '"') | ||
| .replace(/'/g, "'") | ||
| .replace(/'/g, "'") | ||
| .replace(/</g, "<") | ||
| .replace(/>/g, ">") | ||
| .replace(/&/g, "&"); | ||
| } | ||
|
|
||
| // Collect every node that declares an @id, and every bare {"@id"} reference. | ||
| function walk(node, defined, referenced) { | ||
| if (Array.isArray(node)) { | ||
| node.forEach((n) => walk(n, defined, referenced)); | ||
| return; | ||
| } | ||
| if (!node || typeof node !== "object") { | ||
| return; | ||
| } | ||
| const keys = Object.keys(node).filter((k) => k !== "@context"); | ||
| if (node["@id"]) { | ||
| // A node is a *reference* when it carries nothing beyond @id -- including | ||
| // the idiomatic typed form {"@type":"Person","@id":"…"}, which has two keys | ||
| // but still only points at an entity defined elsewhere. Only a node with a | ||
| // real property (name, url, …) *defines* the entity. Treating typed refs as | ||
| // definitions would let a typed pointer at an undefined @id pass silently, | ||
| // which is exactly the dangling case this guard exists to catch. | ||
| const propsBeyondId = keys.filter((k) => k !== "@id" && k !== "@type"); | ||
| if (propsBeyondId.length === 0) { | ||
| referenced.add(node["@id"]); | ||
| } else { | ||
| defined.add(node["@id"]); | ||
| } | ||
| } | ||
| for (const key of keys) { | ||
| walk(node[key], defined, referenced); | ||
| } | ||
| } | ||
|
|
||
| const files = findHtml(buildDir); | ||
|
|
||
| let parseErrors = 0; | ||
| let dangling = 0; | ||
| let blocks = 0; | ||
| let scanned = 0; | ||
| const typeCounts = new Map(); | ||
|
|
||
| for (const file of files) { | ||
| const html = fs.readFileSync(file, "utf8"); | ||
| // Skip noindex pages (archived versions etc.) -- see NOINDEX_RE above. | ||
| if (NOINDEX_RE.test(html)) { | ||
| continue; | ||
| } | ||
| scanned += 1; | ||
| const defined = new Set(); | ||
| const referenced = new Set(); | ||
| const ownUrl = html.match(OG_URL_RE); | ||
| if (ownUrl) { | ||
| defined.add(ownUrl[1]); | ||
| } | ||
| let match; | ||
| SCRIPT_RE.lastIndex = 0; | ||
| while ((match = SCRIPT_RE.exec(html))) { | ||
| blocks += 1; | ||
| let parsed; | ||
| try { | ||
| parsed = JSON.parse(unescapeHtml(match[1])); | ||
| } catch (err) { | ||
| parseErrors += 1; | ||
| console.error(`INVALID JSON-LD ${file}\n ${err.message}`); | ||
| continue; | ||
| } | ||
| const graph = parsed["@graph"] || parsed; | ||
| walk(graph, defined, referenced); | ||
| for (const n of Array.isArray(graph) ? graph : [graph]) { | ||
| const t = n && n["@type"]; | ||
| if (typeof t === "string") { | ||
| typeCounts.set(t, (typeCounts.get(t) || 0) + 1); | ||
| } | ||
| } | ||
| } | ||
| for (const ref of referenced) { | ||
| if (!defined.has(ref)) { | ||
| dangling += 1; | ||
| console.error(`DANGLING @id ${file}\n ${ref}`); | ||
| } | ||
| } | ||
| } | ||
|
|
||
| // The specialization work replaced every generic `Article` with a subtype | ||
| // (TechArticle / APIReference / BlogPosting). Pin that: a stray generic | ||
| // `Article` reappearing is a regression the shape checks above wouldn't catch. | ||
| const genericArticles = typeCounts.get("Article") || 0; | ||
|
|
||
| console.log(`\nPages scanned: ${scanned}`); | ||
| console.log(`JSON-LD blocks: ${blocks}`); | ||
| console.log(`Invalid JSON: ${parseErrors}`); | ||
| console.log(`Dangling @id refs: ${dangling}`); | ||
| console.log(`Generic Article: ${genericArticles}`); | ||
| console.log("\nTop-level @type distribution:"); | ||
| for (const [type, count] of [...typeCounts].sort((a, b) => b[1] - a[1])) { | ||
| console.log(` ${String(count).padStart(5)} ${type}`); | ||
| } | ||
| if (genericArticles) { | ||
| console.error( | ||
| `\nFAIL: ${genericArticles} generic "Article" node(s) -- use TechArticle/APIReference/BlogPosting.` | ||
| ); | ||
| } | ||
|
|
||
| process.exit(parseErrors || dangling || genericArticles ? 1 : 0); |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.