Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -1489,6 +1489,16 @@ follow semantic versioning; release dates are ISO 8601.

### Tests

- **Word holds the DOCX export to its corpus too.** `scripts/docx-visual/word-fidelity.ps1`
runs `DocxFidelityCorpusTest` in two halves around a conversion by a private Word instance:
`export` writes the documents, Word converts them, and `word` measures Word's PDFs against
`word-windows.tsv`. `convert-with-word.ps1` records the SHA-256 of each DOCX it converts, and
a document is measured only when that is the digest of the DOCX this tree exports, so a
conversion left from another tree fails and names the documents rather than passing. Word is
driven from PowerShell, not from the build: Word driven through COM from a process Git Bash
started has stalled on its first document. `convert-with-word.ps1` also repaginates before it
exports, as Word settles pagination on screen.

- **CI holds the DOCX export to its corpus.** The `DOCX Fidelity` job runs
`DocxFidelityCorpusTest` on a pinned Ubuntu image with LibreOffice whenever the engine, a
backend it measures with, the DOCX backend, a template or the corpus changes, against a
Expand Down
5 changes: 4 additions & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -358,7 +358,10 @@ Choose the smallest tests that match the change:
CI runs it on Linux (the `DOCX Fidelity` job) against `libreoffice-linux.tsv`, and uploads
what it measured: LibreOffice sets text a little differently per platform, so a change that
moves documents nearer the page takes that artifact's files as the Linux baseline, and its own
run's as the Windows one.
run's as the Windows one. Microsoft Word, the editor the export answers to, is not in CI: on
Windows with Word, `scripts/docx-visual/word-fidelity.ps1` (`-Update` to rewrite) exports the
corpus, has a private Word instance convert it, and holds it to `word-windows.tsv` — run it
from PowerShell (pwsh), after the install above, before a DOCX export change is opened.

If a change affects public docs, examples, or screenshots, update those assets in the same PR so the repository stays internally consistent.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -26,55 +26,87 @@
* Every template preset's DOCX, set by an editor, stands no further from the engine's PDF than
* its baseline holds it.
*
* <p>The engine draws each corpus document to PDF and exports it to DOCX; LibreOffice converts
* the DOCX to PDF; each document is measured line by line against the page
* <p>The engine draws each corpus document to PDF and exports it to DOCX; an editor converts the
* DOCX to PDF; each document is measured line by line against the page
* ({@link FidelityMeasurement}) and held to the baseline for this editor and operating system
* ({@link FidelityBaseline}). A change to the export that sets any document further from the
* page fails here, wherever the document is.</p>
*
* <p>It runs only when asked for, with {@code -Dgraphcompose.docxFidelity=libreoffice}: it
* needs LibreOffice and takes a few minutes. Asked for and without LibreOffice, it fails rather
* than passing unrun. {@code -Dgraphcompose.docxFidelity.update=true} writes what it measured as
* the baseline, for a change that moves documents nearer the page; the baseline's diff then
* shows the reviewer which documents moved. The documents, their PDFs and the report are left
* in {@code target/docx-fidelity}.</p>
* <p>It runs only when asked for, with {@code -Dgraphcompose.docxFidelity}, and takes a few
* minutes; a value it does not know fails. {@code libreoffice} exports, converts with
* LibreOffice and measures, and fails rather than passing unrun without LibreOffice. Microsoft
* Word, the editor the export answers to, converts outside the build, around two runs of this
* test ({@code scripts/docx-visual/word-fidelity.ps1}): {@code export} writes the documents,
* Word converts them, and {@code word} measures Word's PDFs against {@code word-windows.tsv} —
* each only when the DOCX Word converted, by the SHA-256 its conversion records, is the one
* this tree exports ({@link WordConversion}). {@code -Dgraphcompose.docxFidelity.update=true}
* writes what it measured as the baseline, for a change that moves documents nearer the page;
* the baseline's diff then shows the reviewer which documents moved. The documents, their PDFs
* and the report are left in {@code target/docx-fidelity}.</p>
*/
class DocxFidelityCorpusTest {

private static final String EDITOR = "libreoffice";
private static final String LIBREOFFICE = "libreoffice";
private static final String WORD = "word";
private static final String EXPORT = "export";

@Test
void everyDocumentStandsNoFurtherFromThePageThanItsBaseline() throws Exception {
Assumptions.assumeTrue(EDITOR.equals(System.getProperty("graphcompose.docxFidelity")),
"NOT_RUN: the DOCX fidelity corpus runs with -Dgraphcompose.docxFidelity=libreoffice");
LibreOfficeConverter converter = LibreOfficeConverter.find().orElseThrow(() -> new IllegalStateException(
"the DOCX fidelity corpus was asked for, and no LibreOffice was found: set -Dgraphcompose.soffice"));

String mode = System.getProperty("graphcompose.docxFidelity", "");
Assumptions.assumeTrue(!mode.isEmpty(), "NOT_RUN: the DOCX fidelity corpus runs with "
+ "-Dgraphcompose.docxFidelity=libreoffice, or with Word through scripts/docx-visual/word-fidelity.ps1");
assertThat(mode).as("-Dgraphcompose.docxFidelity").isIn(LIBREOFFICE, WORD, EXPORT);
Path work = Path.of("target", "docx-fidelity").toAbsolutePath();
Path engine = Files.createDirectories(work.resolve("engine"));
Path editor = work.resolve(EDITOR);
// LibreOffice can exit cleanly when a file fails to load: a PDF left from an earlier
// run would then be measured in its place.
deleteTree(editor);
Path engine = work.resolve("engine");
List<DocxCorpusDocument> corpus = corpus();
List<Path> docx = new ArrayList<>();
for (DocxCorpusDocument document : corpus) {
docx.add(export(document, engine));
String note;
Path editor = work.resolve(mode);
if (mode.equals(WORD)) {
// Word converts outside this run (see word-fidelity.ps1), so its PDFs are measured
// only when the DOCX it converted is, to the byte, the one this tree exports.
assertThat(LibreOfficeConverter.platform()).as("Word is measured on Windows").isEqualTo("windows");
WordConversion conversion = WordConversion.read(editor);
Files.createDirectories(engine);
List<String> stale = new ArrayList<>();
for (DocxCorpusDocument document : corpus) {
if (!conversion.convertedFrom(document.stem() + ".docx", exportTo(document, engine))) {
stale.add(document.stem());
}
}
assertThat(stale).as("documents Word converted from other DOCX than this tree exports: run "
+ "scripts/docx-visual/word-fidelity.ps1, which exports and converts first").isEmpty();
note = WORD + " " + conversion.version() + " on " + LibreOfficeConverter.platform();
} else {
// A document no longer in the corpus leaves no DOCX for an editor to convert.
deleteTree(engine);
Files.createDirectories(engine);
List<Path> docx = new ArrayList<>();
for (DocxCorpusDocument document : corpus) {
docx.add(export(document, engine));
}
if (mode.equals(EXPORT)) {
return;
}
LibreOfficeConverter converter = LibreOfficeConverter.find().orElseThrow(() -> new IllegalStateException(
"the DOCX fidelity corpus was asked for, and no LibreOffice was found: set -Dgraphcompose.soffice"));
// LibreOffice can exit cleanly when a file fails to load: a PDF left from an earlier
// run would then be measured in its place.
deleteTree(editor);
converter.convert(docx, editor, work);
note = LIBREOFFICE + " " + converter.build() + " on " + LibreOfficeConverter.platform();
}
converter.convert(docx, editor, work);

List<FidelityMeasurement> measurements = new ArrayList<>();
for (DocxCorpusDocument document : corpus) {
Path converted = editor.resolve(document.stem() + ".pdf");
assertThat(converted).as("LibreOffice's PDF of %s", document.stem()).exists();
assertThat(converted).as("%s's PDF of %s", mode, document.stem()).exists();
measurements.add(FidelityMeasurement.of(document.stem(),
engine.resolve(document.stem() + ".pdf"), converted));
}
String note = EDITOR + " " + converter.build() + " on " + LibreOfficeConverter.platform();
FidelityBaseline.write(work.resolve("measured-" + EDITOR + ".tsv"), note, measurements);
FidelityBaseline.write(work.resolve("measured-" + mode + ".tsv"), note, measurements);

Path baseline = Path.of("src", "test", "resources", "docx-fidelity",
EDITOR + "-" + LibreOfficeConverter.platform() + ".tsv");
mode + "-" + LibreOfficeConverter.platform() + ".tsv");
List<String> regressions = FidelityBaseline.read(baseline).regressions(measurements);
if (Boolean.getBoolean("graphcompose.docxFidelity.update")) {
// Written all the same: what it lets through is the baseline's diff, shown here first.
Expand All @@ -84,7 +116,7 @@ void everyDocumentStandsNoFurtherFromThePageThanItsBaseline() throws Exception {
}
assertThat(regressions)
.as("documents set further from the page than %s holds them (measured: %s)",
baseline, work.resolve("measured-" + EDITOR + ".tsv"))
baseline, work.resolve("measured-" + mode + ".tsv"))
.isEmpty();
}

Expand Down Expand Up @@ -122,6 +154,13 @@ static List<DocxCorpusDocument> corpus() {

/** Writes a document's PDF and its DOCX; returns the DOCX. */
private static Path export(DocxCorpusDocument document, Path dir) throws Exception {
Path docx = dir.resolve(document.stem() + ".docx");
Files.write(docx, exportTo(document, dir));
return docx;
}

/** Writes a document's PDF into {@code dir}; returns its DOCX. */
private static byte[] exportTo(DocxCorpusDocument document, Path dir) throws Exception {
var builder = GraphCompose.document();
if (document.margin() >= 0) {
float margin = (float) document.margin();
Expand All @@ -130,9 +169,7 @@ private static Path export(DocxCorpusDocument document, Path dir) throws Excepti
try (DocumentSession session = builder.create()) {
document.compose().accept(session);
Files.write(dir.resolve(document.stem() + ".pdf"), session.toPdfBytes());
Path docx = dir.resolve(document.stem() + ".docx");
Files.write(docx, session.export(DocxSemanticBackend.builder().deterministic(true).build()));
return docx;
return session.export(DocxSemanticBackend.builder().deterministic(true).build());
}
}
}
Original file line number Diff line number Diff line change
Expand Up @@ -205,6 +205,43 @@ void twoColumnsOnOneBaselineAreTwoLines() throws Exception {
assertThat(PdfLines.of(page).lines()).extracting(PdfLines.Line::key).containsExactly("left", "right");
}

@Test
void aWordConversionIsMeasuredOnlyForTheDocxItConverted() throws Exception {
byte[] docx = docx();
// As convert-with-word.ps1 writes it: UTF-8 with a byte-order mark, the digest in hex.
Files.write(dir.resolve("conversion.json"), ("{\"version\": \"16.0 (16.0.20430)\", \"results\": ["
+ "{\"source\": \"cv-probe.docx\", \"sha256\": \"" + WordConversion.sha256(docx).toUpperCase() + "\"}]}")
.getBytes(java.nio.charset.StandardCharsets.UTF_8));

WordConversion conversion = WordConversion.read(dir);
byte[] edited = docx.clone();
edited[edited.length / 2] ^= 1;

assertThat(conversion.version()).isEqualTo("16.0 (16.0.20430)");
assertThat(conversion.convertedFrom("cv-probe.docx", docx())).as("exported again, the same bytes").isTrue();
assertThat(conversion.convertedFrom("cv-probe.docx", edited)).as("a byte otherwise").isFalse();
assertThat(conversion.convertedFrom("cv-other.docx", docx)).as("a document it never converted").isFalse();
}

@Test
void aWordConversionWithoutItsRecordOrItsBuildIsRefused() throws Exception {
assertThatThrownBy(() -> WordConversion.read(dir)).isInstanceOf(IllegalStateException.class)
.hasMessageContaining("no Word conversion recorded");
Files.writeString(dir.resolve("conversion.json"), "{\"results\": []}");
assertThatThrownBy(() -> WordConversion.read(dir)).isInstanceOf(IllegalStateException.class)
.hasMessageContaining("names no Word build");
}

/** A small document's DOCX, exported as the corpus exports it. */
private static byte[] docx() throws Exception {
try (DocumentSession session = GraphCompose.document().pageSize(400, 300)
.margin(DocumentInsets.of(20)).create()) {
session.pageFlow(flow -> LINES.forEach(flow::addParagraph));
return session.export(com.demcha.compose.document.backend.semantic.docx.DocxSemanticBackend.builder()
.deterministic(true).build());
}
}

@Test
void aMalformedBaselineRowIsRefusedByName() {
assertThatThrownBy(() -> FidelityMeasurement.parse("cv-a\t1\tone\t40\t39\t0.25\t0.5\t1", Map.of()))
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
package com.demcha.compose.document.templates.fidelity;

import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;

import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.security.MessageDigest;
import java.security.NoSuchAlgorithmException;
import java.util.HexFormat;
import java.util.LinkedHashMap;
import java.util.Locale;
import java.util.Map;

/**
* What a Word conversion of the corpus records beside its PDFs, in {@code conversion.json}
* ({@code scripts/docx-visual/convert-with-word.ps1}): the Word build that converted, and the
* SHA-256 of each DOCX it converted.
*
* <p>Word converts outside the build, so its PDFs can be of DOCX another tree exported. A PDF is
* measured only when the DOCX it was converted from is, to the byte, the one this tree exports.</p>
*
* @param version the Word build, as Word states it
* @param converted the SHA-256 of each converted DOCX, in lower-case hex, by its file name
*/
record WordConversion(String version, Map<String, String> converted) {

/** Reads the record a conversion left in {@code dir}; refuses one that is missing or lacks a build. */
static WordConversion read(Path dir) throws IOException {
Path record = dir.resolve("conversion.json");
if (!Files.isRegularFile(record)) {
throw new IllegalStateException("no Word conversion recorded in " + record
+ ": run scripts/docx-visual/word-fidelity.ps1");
}
JsonNode tree = new ObjectMapper().readTree(record.toFile());
String version = tree.path("version").asText("");
if (version.isBlank()) {
throw new IllegalStateException(record + " names no Word build");
}
Map<String, String> converted = new LinkedHashMap<>();
for (JsonNode result : tree.path("results")) {
String source = result.path("source").asText("");
String hash = result.path("sha256").asText("");
if (!source.isBlank() && !hash.isBlank()) {
converted.put(source, hash.toLowerCase(Locale.ROOT));
}
}
return new WordConversion(version, converted);
}

/** Whether Word converted, to the byte, the DOCX given for {@code fileName}. */
boolean convertedFrom(String fileName, byte[] docx) {
return sha256(docx).equals(converted.get(fileName));
}

static String sha256(byte[] bytes) {
try {
return HexFormat.of().formatHex(MessageDigest.getInstance("SHA-256").digest(bytes));
} catch (NoSuchAlgorithmException missing) {
throw new IllegalStateException(missing);
}
}
}
Loading
Loading