Release Schema Tools v0.1.2

This commit is contained in:
2026-09-01 16:20:37 +02:00
parent 2f60a6d0a7
commit f75f2384a8
21 changed files with 1351 additions and 238 deletions
+12
View File
@@ -1,5 +1,17 @@
# Changelog
## 0.1.2 - 2026-09-01
- Refused invalid sample claims when `false` JSON Schemas are reached through local references or required combiners, while selecting the first viable `anyOf` or `oneOf` branch.
- Centralized sample budgets across JSON Schema, OpenAPI, XSD, Relax NG, and Schematron output.
- Applied the 512 KiB aggregate text budget to repeated XML names, attributes, and values, and corrected XML/JSON ceilings to never exceed 2,000 generated values or elements and exactly 20 generated nesting levels.
- Made text truncation UTF-16-safe so it never splits a surrogate pair and corrupts generated XML.
- Preserved direct-root Relax NG `element` patterns in generated samples.
- Added a monotonic 50,000-step work ceiling that refuses over-budget generation, so repeated references, failed alternatives, and non-emitting pattern fan-out cannot amplify work or bypass later mandatory constraints.
- Replaced costly live DOM child-collection conversions with linear sibling-pointer walks and cached bounded XSD type shapes.
- Cached reused schema/reference and XML pattern lookups, and omitted duplicate derived XML attributes so generated samples remain well formed.
- Added adversarial regressions for referenced boolean schemas, exhausted combiners and required values, amplified shared XSD values, wide XML containers, repeated Relax NG references, and exact JSON/XML node limits.
## 0.1.1 - 2026-09-01
- Bounded literal values, aggregate text, derived XML names, collections, and final output during sample generation.
+2 -2
View File
@@ -14,7 +14,7 @@ Schema Tools is a standalone local-first application in the [add·ideas Toolbox]
- Conservative comparisons for required properties, types, enums, declarations, operations, parameters, and responses
- Inert rendering, offline PWA support, responsive Toolbox shell integration, and deterministic release archives
Each document is limited to 2 MiB of text, the workspace to 8 MiB and 20 documents, and parsed trees, references, instances, samples, validation work, and diagnostic output have independent bounds. Generated samples are additionally limited to 2 MiB of serialized output, 512 KiB of retained text, and 1,024 characters per derived literal. DTD/entity declarations, remote references, Schematron XPath, extension code, JSON Schema `pattern`, and `patternProperties` are never executed. JSON Schema validation covers a documented assertion subset rather than claiming full specification conformance. XML languages receive structural inspection—not instance validation—and OpenAPI checks are intentionally focused.
Each document is limited to 2 MiB of text, the workspace to 8 MiB and 20 documents, and parsed trees, references, instances, samples, validation work, and diagnostic output have independent bounds. One shared generator budget limits samples to 50,000 monotonic work steps, 2,000 generated JSON values or XML elements, exactly 20 generated nesting levels, 512 KiB of aggregate derived text, 1,024 UTF-16 code units per literal without splitting surrogate pairs, and 2 MiB of serialized output. Generation is refused if another work step would exceed that ceiling; structural and text ceilings instead return a visibly bounded heuristic sample where one can still be formed. DTD/entity declarations, remote references, Schematron XPath, extension code, JSON Schema `pattern`, and `patternProperties` are never executed. JSON Schema validation covers a documented assertion subset rather than claiming full specification conformance. XML languages receive structural inspection—not instance validation—and OpenAPI checks are intentionally focused.
See [Architecture](docs/ARCHITECTURE.md) and [Privacy and security](docs/PRIVACY-SECURITY.md) for the exact capability boundaries.
@@ -30,7 +30,7 @@ npm run test:browser
## Release
`npm run release:artifact` creates deterministic `release/schema-tools-0.1.1.zip` and checksum files.
`npm run release:artifact` creates deterministic `release/schema-tools-0.1.2.zip` and checksum files.
## Licence
+2 -2
View File
@@ -1,7 +1,7 @@
# Corresponding source
The corresponding source for Schema Tools 0.1.1 is available at:
The corresponding source for Schema Tools 0.1.2 is available at:
https://git.add-ideas.de/lotobo/schema-tools/src/tag/v0.1.1
https://git.add-ideas.de/lotobo/schema-tools/src/tag/v0.1.2
Build with Node.js 22, npm 11, `npm ci`, and `npm run release:artifact`.
+1 -1
View File
@@ -1,6 +1,6 @@
# Third-party notices
Schema Tools 0.1.1 directly depends on these runtime packages:
Schema Tools 0.1.2 directly depends on these runtime packages:
| Package | Pinned version | Declared licence |
| -------------------------------- | -------------: | ---------------- |
+4 -2
View File
@@ -6,6 +6,8 @@ JSON and YAML values are converted into an acyclic, prototype-safe JSON model wi
Every reference is classified before validation. Only fragments and relative filenames supplied in the current workspace are eligible. Root `$id`/legacy `id` values are deliberately ignored so relative references remain anchored to workspace filenames; nested identifier scopes, named anchors, dynamic/recursive references, and unevaluated keywords are outside the subset and cause a visible refusal. All participating JSON Schema files must use one declared draft. Formats and unknown extension keywords are annotations. Schemas containing `pattern` or `patternProperties` are refused because JavaScript regular-expression execution cannot be reliably time-bounded; they remain inspectable. Decimal arithmetic uses JavaScript numbers, so `multipleOf` applies a small floating-point tolerance.
OpenAPI JSON/YAML receives focused document, operation, response, reference, sample, and comparison logic. It does not run requests and does not claim full OpenAPI conformance. XML uses the browser's inert `DOMParser` only after rejecting DTD and entity declarations. XSD, Relax NG XML syntax, and Schematron are checked for well-formedness and structurally inventoried. Schematron XPath and extensions are retained as text and never executed. XSD and Relax NG sample generation deliberately follows a bounded first branch and is labelled heuristic. Across languages, sample generation copies literals rather than retaining source objects, caps individual derived values and names, limits aggregate retained text and nodes, and rejects serialized output above 2 MiB.
OpenAPI JSON/YAML receives focused document, operation, response, reference, sample, and comparison logic. It does not run requests and does not claim full OpenAPI conformance. XML uses the browser's inert `DOMParser` only after rejecting DTD and entity declarations. Element counting and depth inspection use a linear sibling-pointer walk, avoiding repeated conversion of live DOM child collections. XSD, Relax NG XML syntax, and Schematron are checked for well-formedness and structurally inventoried. Schematron XPath and extensions are retained as text and never executed. XSD and Relax NG sample generation deliberately follows a bounded first branch and is labelled heuristic. XSD named-type shapes are inspected once and cached within one generation; Relax NG grammars and direct-root `element` patterns use the same renderer, and pattern-only wrappers do not consume generated nesting depth.
The PWA uses only relative URLs, so the same build works standalone or below a nested portal route. Its service worker caches same-origin files from its own scope. No worker, storage database, server API, telemetry path, or remote schema resolver exists in v0.1.1.
All sample paths share one budget implementation. It admits at most 50,000 monotonic generator work steps and 2,000 generated JSON values or XML elements across exactly 20 generated levels (the root is level one), accounts for at most 512 KiB of aggregate derived keys, names, attribute values, and content, caps each copied literal at 1,024 UTF-16 code units without splitting a surrogate pair, and rejects serialized output above 2 MiB. References and schema/pattern combiners do not consume generated depth, but every build, copy, reference, alternative, XML type inspection, and XML pattern visit consumes work before expansion. Per-generation caches ensure repeated JSON references, wide property collections, XML child collections, inline types, text, and name normalization are not rescanned without bound. The work counter is deliberately not restored when a failed heuristic alternative rolls back its node/text checkpoint, and generation is refused rather than returning an ambiguously partial result when a 50,001st work step is attempted. Node, depth, and text ceilings can instead omit bounded material with a visible notice. Repeated references consume node and text counters for every emitted occurrence. Duplicate or fallback-colliding derived XML attribute names are omitted with a notice to preserve well-formed output. A `false` JSON Schema reached through a selected local reference or mandatory `allOf` branch aborts generation; `anyOf` and `oneOf` skip definitely impossible boolean branches and use the first viable heuristic branch only while the complete attempt remains within budget.
The PWA uses only relative URLs, so the same build works standalone or below a nested portal route. Its service worker caches same-origin files from its own scope. No worker, storage database, server API, telemetry path, or remote schema resolver exists in v0.1.2.
+2 -2
View File
@@ -2,8 +2,8 @@
Schema sources, instances, diagnostics, comparisons, and generated samples stay in page memory. There are no accounts, analytics, telemetry, persistence, remote imports, or runtime third-party assets. Explicit source links are normal navigation only. Clearing or closing the page releases application references but cannot promise forensic erasure from browser or operating-system memory.
Each source is limited to 2 MiB of text; a workspace is limited to 20 documents and 8 MiB. JSON/YAML trees are capped at 25,000 values, 48 levels, and 2,000 entries in one collection. XML is capped at 25,000 elements and 48 levels. References, instances, and diagnostics have separate caps. Sample generation is limited to 2,000 nodes, 20 levels, 512 KiB of retained text, 1,024 characters per copied literal, and 2 MiB of serialized output. File byte gates are deliberately conservative before `File.text()` decoding.
Each source is limited to 2 MiB of text; a workspace is limited to 20 documents and 8 MiB. JSON/YAML trees are capped at 25,000 values, 48 levels, and 2,000 entries in one collection. XML is capped at 25,000 elements and 48 levels. References, instances, and diagnostics have separate caps. A shared sample budget is limited to 50,000 monotonic work steps, exactly 2,000 generated JSON values or XML elements, exactly 20 generated nesting levels, 512 KiB of aggregate derived text, 1,024 UTF-16 code units per copied literal without splitting surrogate pairs, and 2 MiB of serialized output. Reusing a local definition consumes the same aggregate counters on every attempt and generated occurrence; failed heuristic alternatives never refund work. Reaching the work ceiling refuses the generation operation, so exhaustion cannot be mistaken for a viable choice or skip a later mandatory constraint. File byte gates are deliberately conservative before `File.text()` decoding.
Prototype-sensitive JSON keys, cyclic YAML aliases, NUL input, DTD/entity declarations, absolute/remote/escaping references, and executable schema extensions are rejected. Schematron expressions and imported XML are never executed. JSON Schema formats are annotations, and regex-bearing schema keywords are not executed. The focused validator interprets eligible schemas without dynamic code generation, remote loading, custom code, or `unsafe-eval`; work and diagnostic counts are capped.
Sample generation and compatibility results are review aids. A generated document is not guaranteed to satisfy every constraint, and an absence of reported changes does not prove compatibility. JSON validation requires a single declared draft and uses local workspace filenames rather than identifier URIs. Named anchors, nested identifier scopes, dynamic/recursive references, unevaluated keywords, regular expressions, and full meta-schema validation are outside v0.1.1. XSD, Relax NG, and Schematron instance validation is also outside v0.1.1; OpenAPI inspection is not a full conformance certification.
Sample generation and compatibility results are review aids. A generated document is not guaranteed to satisfy every constraint, and an absence of reported changes does not prove compatibility. Generation refuses a definitely impossible selected `false` JSON Schema branch but does not prove broader satisfiability. JSON validation requires a single declared draft and uses local workspace filenames rather than identifier URIs. Named anchors, nested identifier scopes, dynamic/recursive references, unevaluated keywords, regular expressions, and full meta-schema validation are outside v0.1.2. XSD, Relax NG, and Schematron instance validation is also outside v0.1.2; OpenAPI inspection is not a full conformance certification.
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "schema-tools",
"version": "0.1.1",
"version": "0.1.2",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "schema-tools",
"version": "0.1.1",
"version": "0.1.2",
"license": "GPL-3.0-or-later",
"dependencies": {
"@add-ideas/toolbox-contract": "0.2.3",
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "schema-tools",
"version": "0.1.1",
"version": "0.1.2",
"description": "Inspect, validate, compare, and derive examples from schemas locally in the browser.",
"license": "GPL-3.0-or-later",
"author": "Albrecht Degering",
+12
View File
@@ -1,5 +1,17 @@
# Changelog
## 0.1.2 - 2026-09-01
- Refused invalid sample claims when `false` JSON Schemas are reached through local references or required combiners, while selecting the first viable `anyOf` or `oneOf` branch.
- Centralized sample budgets across JSON Schema, OpenAPI, XSD, Relax NG, and Schematron output.
- Applied the 512 KiB aggregate text budget to repeated XML names, attributes, and values, and corrected XML/JSON ceilings to never exceed 2,000 generated values or elements and exactly 20 generated nesting levels.
- Made text truncation UTF-16-safe so it never splits a surrogate pair and corrupts generated XML.
- Preserved direct-root Relax NG `element` patterns in generated samples.
- Added a monotonic 50,000-step work ceiling that refuses over-budget generation, so repeated references, failed alternatives, and non-emitting pattern fan-out cannot amplify work or bypass later mandatory constraints.
- Replaced costly live DOM child-collection conversions with linear sibling-pointer walks and cached bounded XSD type shapes.
- Cached reused schema/reference and XML pattern lookups, and omitted duplicate derived XML attributes so generated samples remain well formed.
- Added adversarial regressions for referenced boolean schemas, exhausted combiners and required values, amplified shared XSD values, wide XML containers, repeated Relax NG references, and exact JSON/XML node limits.
## 0.1.1 - 2026-09-01
- Bounded literal values, aggregate text, derived XML names, collections, and final output during sample generation.
+2 -2
View File
@@ -14,7 +14,7 @@ Schema Tools is a standalone local-first application in the [add·ideas Toolbox]
- Conservative comparisons for required properties, types, enums, declarations, operations, parameters, and responses
- Inert rendering, offline PWA support, responsive Toolbox shell integration, and deterministic release archives
Each document is limited to 2 MiB of text, the workspace to 8 MiB and 20 documents, and parsed trees, references, instances, samples, validation work, and diagnostic output have independent bounds. Generated samples are additionally limited to 2 MiB of serialized output, 512 KiB of retained text, and 1,024 characters per derived literal. DTD/entity declarations, remote references, Schematron XPath, extension code, JSON Schema `pattern`, and `patternProperties` are never executed. JSON Schema validation covers a documented assertion subset rather than claiming full specification conformance. XML languages receive structural inspection—not instance validation—and OpenAPI checks are intentionally focused.
Each document is limited to 2 MiB of text, the workspace to 8 MiB and 20 documents, and parsed trees, references, instances, samples, validation work, and diagnostic output have independent bounds. One shared generator budget limits samples to 50,000 monotonic work steps, 2,000 generated JSON values or XML elements, exactly 20 generated nesting levels, 512 KiB of aggregate derived text, 1,024 UTF-16 code units per literal without splitting surrogate pairs, and 2 MiB of serialized output. Generation is refused if another work step would exceed that ceiling; structural and text ceilings instead return a visibly bounded heuristic sample where one can still be formed. DTD/entity declarations, remote references, Schematron XPath, extension code, JSON Schema `pattern`, and `patternProperties` are never executed. JSON Schema validation covers a documented assertion subset rather than claiming full specification conformance. XML languages receive structural inspection—not instance validation—and OpenAPI checks are intentionally focused.
See [Architecture](docs/ARCHITECTURE.md) and [Privacy and security](docs/PRIVACY-SECURITY.md) for the exact capability boundaries.
@@ -30,7 +30,7 @@ npm run test:browser
## Release
`npm run release:artifact` creates deterministic `release/schema-tools-0.1.1.zip` and checksum files.
`npm run release:artifact` creates deterministic `release/schema-tools-0.1.2.zip` and checksum files.
## Licence
+2 -2
View File
@@ -1,7 +1,7 @@
# Corresponding source
The corresponding source for Schema Tools 0.1.1 is available at:
The corresponding source for Schema Tools 0.1.2 is available at:
https://git.add-ideas.de/lotobo/schema-tools/src/tag/v0.1.1
https://git.add-ideas.de/lotobo/schema-tools/src/tag/v0.1.2
Build with Node.js 22, npm 11, `npm ci`, and `npm run release:artifact`.
+1 -1
View File
@@ -1,6 +1,6 @@
# Third-party notices
Schema Tools 0.1.1 directly depends on these runtime packages:
Schema Tools 0.1.2 directly depends on these runtime packages:
| Package | Pinned version | Declared licence |
| -------------------------------- | -------------: | ---------------- |
+4 -2
View File
@@ -6,6 +6,8 @@ JSON and YAML values are converted into an acyclic, prototype-safe JSON model wi
Every reference is classified before validation. Only fragments and relative filenames supplied in the current workspace are eligible. Root `$id`/legacy `id` values are deliberately ignored so relative references remain anchored to workspace filenames; nested identifier scopes, named anchors, dynamic/recursive references, and unevaluated keywords are outside the subset and cause a visible refusal. All participating JSON Schema files must use one declared draft. Formats and unknown extension keywords are annotations. Schemas containing `pattern` or `patternProperties` are refused because JavaScript regular-expression execution cannot be reliably time-bounded; they remain inspectable. Decimal arithmetic uses JavaScript numbers, so `multipleOf` applies a small floating-point tolerance.
OpenAPI JSON/YAML receives focused document, operation, response, reference, sample, and comparison logic. It does not run requests and does not claim full OpenAPI conformance. XML uses the browser's inert `DOMParser` only after rejecting DTD and entity declarations. XSD, Relax NG XML syntax, and Schematron are checked for well-formedness and structurally inventoried. Schematron XPath and extensions are retained as text and never executed. XSD and Relax NG sample generation deliberately follows a bounded first branch and is labelled heuristic. Across languages, sample generation copies literals rather than retaining source objects, caps individual derived values and names, limits aggregate retained text and nodes, and rejects serialized output above 2 MiB.
OpenAPI JSON/YAML receives focused document, operation, response, reference, sample, and comparison logic. It does not run requests and does not claim full OpenAPI conformance. XML uses the browser's inert `DOMParser` only after rejecting DTD and entity declarations. Element counting and depth inspection use a linear sibling-pointer walk, avoiding repeated conversion of live DOM child collections. XSD, Relax NG XML syntax, and Schematron are checked for well-formedness and structurally inventoried. Schematron XPath and extensions are retained as text and never executed. XSD and Relax NG sample generation deliberately follows a bounded first branch and is labelled heuristic. XSD named-type shapes are inspected once and cached within one generation; Relax NG grammars and direct-root `element` patterns use the same renderer, and pattern-only wrappers do not consume generated nesting depth.
The PWA uses only relative URLs, so the same build works standalone or below a nested portal route. Its service worker caches same-origin files from its own scope. No worker, storage database, server API, telemetry path, or remote schema resolver exists in v0.1.1.
All sample paths share one budget implementation. It admits at most 50,000 monotonic generator work steps and 2,000 generated JSON values or XML elements across exactly 20 generated levels (the root is level one), accounts for at most 512 KiB of aggregate derived keys, names, attribute values, and content, caps each copied literal at 1,024 UTF-16 code units without splitting a surrogate pair, and rejects serialized output above 2 MiB. References and schema/pattern combiners do not consume generated depth, but every build, copy, reference, alternative, XML type inspection, and XML pattern visit consumes work before expansion. Per-generation caches ensure repeated JSON references, wide property collections, XML child collections, inline types, text, and name normalization are not rescanned without bound. The work counter is deliberately not restored when a failed heuristic alternative rolls back its node/text checkpoint, and generation is refused rather than returning an ambiguously partial result when a 50,001st work step is attempted. Node, depth, and text ceilings can instead omit bounded material with a visible notice. Repeated references consume node and text counters for every emitted occurrence. Duplicate or fallback-colliding derived XML attribute names are omitted with a notice to preserve well-formed output. A `false` JSON Schema reached through a selected local reference or mandatory `allOf` branch aborts generation; `anyOf` and `oneOf` skip definitely impossible boolean branches and use the first viable heuristic branch only while the complete attempt remains within budget.
The PWA uses only relative URLs, so the same build works standalone or below a nested portal route. Its service worker caches same-origin files from its own scope. No worker, storage database, server API, telemetry path, or remote schema resolver exists in v0.1.2.
+2 -2
View File
@@ -2,8 +2,8 @@
Schema sources, instances, diagnostics, comparisons, and generated samples stay in page memory. There are no accounts, analytics, telemetry, persistence, remote imports, or runtime third-party assets. Explicit source links are normal navigation only. Clearing or closing the page releases application references but cannot promise forensic erasure from browser or operating-system memory.
Each source is limited to 2 MiB of text; a workspace is limited to 20 documents and 8 MiB. JSON/YAML trees are capped at 25,000 values, 48 levels, and 2,000 entries in one collection. XML is capped at 25,000 elements and 48 levels. References, instances, and diagnostics have separate caps. Sample generation is limited to 2,000 nodes, 20 levels, 512 KiB of retained text, 1,024 characters per copied literal, and 2 MiB of serialized output. File byte gates are deliberately conservative before `File.text()` decoding.
Each source is limited to 2 MiB of text; a workspace is limited to 20 documents and 8 MiB. JSON/YAML trees are capped at 25,000 values, 48 levels, and 2,000 entries in one collection. XML is capped at 25,000 elements and 48 levels. References, instances, and diagnostics have separate caps. A shared sample budget is limited to 50,000 monotonic work steps, exactly 2,000 generated JSON values or XML elements, exactly 20 generated nesting levels, 512 KiB of aggregate derived text, 1,024 UTF-16 code units per copied literal without splitting surrogate pairs, and 2 MiB of serialized output. Reusing a local definition consumes the same aggregate counters on every attempt and generated occurrence; failed heuristic alternatives never refund work. Reaching the work ceiling refuses the generation operation, so exhaustion cannot be mistaken for a viable choice or skip a later mandatory constraint. File byte gates are deliberately conservative before `File.text()` decoding.
Prototype-sensitive JSON keys, cyclic YAML aliases, NUL input, DTD/entity declarations, absolute/remote/escaping references, and executable schema extensions are rejected. Schematron expressions and imported XML are never executed. JSON Schema formats are annotations, and regex-bearing schema keywords are not executed. The focused validator interprets eligible schemas without dynamic code generation, remote loading, custom code, or `unsafe-eval`; work and diagnostic counts are capped.
Sample generation and compatibility results are review aids. A generated document is not guaranteed to satisfy every constraint, and an absence of reported changes does not prove compatibility. JSON validation requires a single declared draft and uses local workspace filenames rather than identifier URIs. Named anchors, nested identifier scopes, dynamic/recursive references, unevaluated keywords, regular expressions, and full meta-schema validation are outside v0.1.1. XSD, Relax NG, and Schematron instance validation is also outside v0.1.1; OpenAPI inspection is not a full conformance certification.
Sample generation and compatibility results are review aids. A generated document is not guaranteed to satisfy every constraint, and an absence of reported changes does not prove compatibility. Generation refuses a definitely impossible selected `false` JSON Schema branch but does not prove broader satisfiability. JSON validation requires a single declared draft and uses local workspace filenames rather than identifier URIs. Named anchors, nested identifier scopes, dynamic/recursive references, unevaluated keywords, regular expressions, and full meta-schema validation are outside v0.1.2. XSD, Relax NG, and Schematron instance validation is also outside v0.1.2; OpenAPI inspection is not a full conformance certification.
+1 -1
View File
@@ -1,5 +1,5 @@
const CACHE_PREFIX = "schema-tools-shell-";
const CACHE_NAME = CACHE_PREFIX + "0.1.1";
const CACHE_NAME = CACHE_PREFIX + "0.1.2";
const CORE = ["./", "./manifest.webmanifest", "./favicon.svg"];
self.addEventListener("install", (event) => {
event.waitUntil(
+1 -1
View File
@@ -3,7 +3,7 @@
"schemaVersion": 1,
"id": "de.add-ideas.schema-tools",
"name": "Schema Tools",
"version": "0.1.1",
"version": "0.1.2",
"description": "Inspect, validate, compare, and derive schema examples locally.",
"entry": "./",
"icon": "./favicon.svg",
+637 -202
View File
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -3,7 +3,7 @@
"schemaVersion": 1,
"id": "de.add-ideas.schema-tools",
"name": "Schema Tools",
"version": "0.1.1",
"version": "0.1.2",
"description": "Inspect, validate, compare, and derive schema examples locally.",
"entry": "./",
"icon": "./favicon.svg",
+1 -1
View File
@@ -1 +1 @@
export const APP_VERSION = "0.1.1";
export const APP_VERSION = "0.1.2";
+1 -1
View File
@@ -70,7 +70,7 @@ test("serves release identity and hardened headers", async ({ request }) => {
const manifest = await request.get("/deep/nested/schema/toolbox-app.json");
await expect(manifest.json()).resolves.toMatchObject({
id: "de.add-ideas.schema-tools",
version: "0.1.1",
version: "0.1.2",
entry: "./",
privacy: { processing: "local", telemetry: false },
});
+650
View File
@@ -37,6 +37,77 @@ const support = {
}),
};
function nestedJsonSchema(levels: number): object {
let schema: object = { const: "leaf" };
for (let level = 1; level < levels; level += 1)
schema = {
type: "object",
required: ["child"],
properties: { child: schema },
};
return schema;
}
function jsonValueDepth(value: unknown): number {
if (!value || typeof value !== "object") return 1;
const children = Object.values(value);
return 1 + (children.length ? Math.max(...children.map(jsonValueDepth)) : 0);
}
function xsdChain(levels: number): string {
const types = Array.from({ length: levels - 1 }, (_, index) => {
const number = index + 1;
const childType = number === levels - 1 ? "xs:string" : `Type${number + 1}`;
return `<xs:complexType name="Type${number}"><xs:sequence><xs:element name="level${number + 1}" type="${childType}"/></xs:sequence></xs:complexType>`;
}).join("");
return `<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema">${types}<xs:element name="level1" type="Type1"/></xs:schema>`;
}
function relaxNgChain(levels: number): string {
let pattern = "<text/>";
for (let level = levels; level >= 1; level -= 1)
pattern = `<element name="level${level}"><group>${pattern}</group></element>`;
return `<grammar xmlns="http://relaxng.org/ns/structure/1.0"><start>${pattern}</start></grammar>`;
}
function hasUnpairedSurrogate(value: string): boolean {
for (let index = 0; index < value.length; index += 1) {
const code = value.charCodeAt(index);
if (code >= 0xd800 && code <= 0xdbff) {
const next = value.charCodeAt(index + 1);
if (next < 0xdc00 || next > 0xdfff) return true;
index += 1;
} else if (code >= 0xdc00 && code <= 0xdfff) return true;
}
return false;
}
const amplifiedChoiceSupport = {
name: "choices.json",
source: JSON.stringify({
$defs: {
impossible: {
anyOf: Array.from({ length: 600 }, () => false),
},
amplified: {
anyOf: Array.from({ length: 100 }, () => ({
$ref: "#/$defs/impossible",
})),
},
},
}),
};
function captureError(run: () => unknown): Error {
try {
run();
} catch (error) {
if (error instanceof Error) return error;
throw error;
}
throw new Error("Expected operation to throw.");
}
describe("schema workspace", () => {
it("resolves local references, generates a sample, and validates instances", () => {
const workspace = inspectWorkspace([entry, support], entry.name);
@@ -99,6 +170,226 @@ describe("schema workspace", () => {
expect(sample.notices.join("\n")).toMatch(/truncated.*safety bound/iu);
});
it("applies the exact JSON node budget to copied literal collections", () => {
const workspace = inspectWorkspace([
{
name: "literal-array.json",
source: JSON.stringify({
default: Array.from({ length: 100 }, () =>
Array.from({ length: 20 }, () => null),
),
}),
},
]);
const sample = generateSample(workspace);
const value = JSON.parse(sample.output) as unknown;
const countNodes = (item: unknown): number =>
item && typeof item === "object"
? 1 +
Object.values(item).reduce((sum, child) => sum + countNodes(child), 0)
: 1;
expect(countNodes(value)).toBe(SCHEMA_LIMITS.sampleNodes);
expect(sample.notices.join("\n")).toMatch(/node.*safety bound/iu);
});
it("monotonically bounds amplified JSON choice attempts through reused refs", () => {
const properties = Object.fromEntries(
Array.from({ length: 100 }, (_, index) => [
`value${index}`,
{ $ref: "choices.json#/$defs/value" },
]),
);
const workspace = inspectWorkspace([
{
name: "root.json",
source: JSON.stringify({
type: "object",
required: Object.keys(properties),
properties,
}),
},
{
name: "choices.json",
source: JSON.stringify({
$defs: {
value: {
anyOf: [
...Array.from({ length: 600 }, () => false),
{ const: "viable" },
],
},
},
}),
},
]);
const started = performance.now();
const error = captureError(() => generateSample(workspace));
const elapsed = performance.now() - started;
expect(error.message).toMatch(/sample budget was exhausted/iu);
expect(error.message).not.toMatch(/first viable/iu);
expect(elapsed).toBeLessThan(2_000);
});
it("never treats exhausted all-impossible anyOf or oneOf branches as viable", () => {
for (const keyword of ["anyOf", "oneOf"] as const) {
const workspace = inspectWorkspace([
{
name: `${keyword}.json`,
source: JSON.stringify({
[keyword]: Array.from({ length: 100 }, () => ({
$ref: "choices.json#/$defs/impossible",
})),
}),
},
amplifiedChoiceSupport,
]);
const error = captureError(() => generateSample(workspace));
expect(error.message).toMatch(/sample budget was exhausted/iu);
expect(error.message).not.toMatch(/first viable/iu);
}
});
it("refuses a partial mandatory allOf when prior work exhausts the budget", () => {
const properties = Object.fromEntries(
Array.from({ length: 100 }, (_, index) => [
`value${index}`,
{ $ref: "choices.json#/$defs/amplified" },
]),
);
const workspace = inspectWorkspace([
{
name: "all-of.json",
source: JSON.stringify({
allOf: [
{
type: "object",
required: Object.keys(properties),
properties,
},
false,
],
}),
},
amplifiedChoiceSupport,
]);
expect(() => generateSample(workspace)).toThrow(
/sample budget was exhausted/iu,
);
});
it("propagates optional-property exhaustion before later required constraints", () => {
for (const requiredSchema of [{ const: "required" }, false] as const) {
const workspace = inspectWorkspace([
{
name: "optional-before-required.json",
source: JSON.stringify({
type: "object",
required: ["required"],
properties: {
optional: { $ref: "choices.json#/$defs/amplified" },
required: requiredSchema,
},
}),
},
amplifiedChoiceSupport,
]);
expect(() => generateSample(workspace)).toThrow(
/sample budget was exhausted/iu,
);
}
});
it("propagates work exhaustion from a required array item", () => {
const workspace = inspectWorkspace([
{
name: "required-item.json",
source: JSON.stringify({
type: "array",
minItems: 1,
items: { $ref: "choices.json#/$defs/amplified" },
}),
},
amplifiedChoiceSupport,
]);
expect(() => generateSample(workspace)).toThrow(
/sample budget was exhausted/iu,
);
});
it("charges cached wide JSON property visits on every reused schema", () => {
const wideProperties = Object.fromEntries(
Array.from({ length: 2_000 }, (_, index) => [
`optional${index}`,
{ type: "string" },
]),
);
const workspace = inspectWorkspace([
{
name: "wide-references.json",
source: JSON.stringify({
allOf: Array.from({ length: 2_000 }, () => ({
$ref: "wide.json#/$defs/wide",
})),
}),
},
{
name: "wide.json",
source: JSON.stringify({
$defs: {
wide: { type: "object", properties: wideProperties },
},
}),
},
]);
const started = performance.now();
const error = captureError(() => generateSample(workspace));
expect(error.message).toMatch(/sample budget was exhausted/iu);
expect(performance.now() - started).toBeLessThan(2_000);
});
it("admits exactly 20 generated JSON levels and omits level 21", () => {
const exact = inspectWorkspace([
{
name: "depth-20.json",
source: JSON.stringify(nestedJsonSchema(20)),
},
]);
const exactSample = generateSample(exact);
expect(jsonValueDepth(JSON.parse(exactSample.output))).toBe(20);
expect(exactSample.notices.join("\n")).not.toMatch(/level safety bound/iu);
const excessive = inspectWorkspace([
{
name: "depth-21.json",
source: JSON.stringify(nestedJsonSchema(21)),
},
]);
const boundedSample = generateSample(excessive);
expect(jsonValueDepth(JSON.parse(boundedSample.output))).toBe(20);
expect(boundedSample.notices.join("\n")).toMatch(/level safety bound/iu);
});
it("truncates JSON literals without splitting surrogate pairs", () => {
const value = `a${"😀".repeat(600)}`;
const workspace = inspectWorkspace([
{
name: "unicode.json",
source: JSON.stringify({ default: value }),
},
]);
const generated = JSON.parse(generateSample(workspace).output) as string;
expect(generated.length).toBeLessThanOrEqual(
SCHEMA_LIMITS.sampleValueChars,
);
expect(hasUnpairedSurrogate(generated)).toBe(false);
});
it("refuses to claim a sample for the always-invalid false schema", () => {
const workspace = inspectWorkspace([
{ name: "impossible.json", source: "false" },
@@ -106,6 +397,34 @@ describe("schema workspace", () => {
expect(() => generateSample(workspace)).toThrow(/no valid sample/iu);
});
it("refuses false schemas reached through references and required combiners", () => {
const referenced = inspectWorkspace([
{ name: "root.json", source: JSON.stringify({ $ref: "defs.json" }) },
{ name: "defs.json", source: "false" },
]);
expect(() => generateSample(referenced)).toThrow(/no valid sample/iu);
const combined = inspectWorkspace([
{
name: "combined.json",
source: JSON.stringify({ allOf: [{ type: "object" }, false] }),
},
]);
expect(() => generateSample(combined)).toThrow(/no valid sample/iu);
});
it("skips impossible alternatives when a viable anyOf sample exists", () => {
const workspace = inspectWorkspace([
{
name: "choice.json",
source: JSON.stringify({
anyOf: [false, { type: "string", const: "viable" }],
}),
},
]);
expect(JSON.parse(generateSample(workspace).output)).toBe("viable");
});
it("blocks remote and missing references without attempting resolution", () => {
const workspace = inspectWorkspace([
{
@@ -261,6 +580,154 @@ describe("schema languages", () => {
expect(sample.notices[0]).toContain("Heuristic XSD");
});
it("applies the aggregate text budget to repeated XSD type literals", () => {
const branches = Array.from(
{ length: 50 },
(_, index) => `<xs:element name="branch${index}" type="Branch"/>`,
).join("");
const leaves = Array.from(
{ length: 50 },
(_, index) => `<xs:element name="item${index}" type="LargeText"/>`,
).join("");
const workspace = inspectWorkspace([
{
name: "amplified.xsd",
source: `<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"><xs:simpleType name="LargeText"><xs:restriction base="xs:string"><xs:enumeration value="${"x".repeat(SCHEMA_LIMITS.sampleValueChars)}"/></xs:restriction></xs:simpleType><xs:complexType name="Branch"><xs:sequence>${leaves}</xs:sequence></xs:complexType><xs:element name="root"><xs:complexType><xs:sequence>${branches}</xs:sequence></xs:complexType></xs:element></xs:schema>`,
},
]);
const sample = generateSample(workspace);
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(parsed.documentElement.textContent?.length ?? 0).toBeLessThanOrEqual(
SCHEMA_LIMITS.sampleTextChars,
);
expect(parsed.querySelectorAll("*").length).toBeLessThanOrEqual(
SCHEMA_LIMITS.sampleNodes,
);
expect(sample.output.length).toBeLessThanOrEqual(
SCHEMA_LIMITS.sampleOutputChars,
);
expect(sample.notices.join("\n")).toMatch(/text safety bound/iu);
});
it("linearly inspects and preindexes a large reused XSD type", () => {
const leaves = Array.from(
{ length: 20_000 },
() => `<xs:element name="leaf" type="xs:string"/>`,
).join("");
const branches = Array.from(
{ length: 50 },
(_, index) => `<xs:element name="branch${index}" type="Heavy"/>`,
).join("");
const source = `<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"><xs:complexType name="Heavy"><xs:sequence>${leaves}</xs:sequence></xs:complexType><xs:complexType name="Root"><xs:sequence>${branches}</xs:sequence></xs:complexType><xs:element name="root" type="Root"/></xs:schema>`;
const inspectStarted = performance.now();
const workspace = inspectWorkspace([{ name: "large-flat.xsd", source }]);
const inspectElapsed = performance.now() - inspectStarted;
const generateStarted = performance.now();
const sample = generateSample(workspace);
const generateElapsed = performance.now() - generateStarted;
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(parsed.querySelectorAll("*").length).toBeLessThanOrEqual(
SCHEMA_LIMITS.sampleNodes,
);
expect(inspectElapsed).toBeLessThan(2_000);
expect(generateElapsed).toBeLessThan(2_000);
});
it("linearly inspects a wide XSD schema container", () => {
const annotations = Array.from(
{ length: 20_000 },
() => "<xs:annotation/>",
).join("");
const source = `<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"><xs:element name="root"/>${annotations}</xs:schema>`;
const started = performance.now();
const workspace = inspectWorkspace([{ name: "wide-root.xsd", source }]);
const elapsed = performance.now() - started;
expect(workspace.documents.get(workspace.entry)?.language).toBe("xsd");
expect(elapsed).toBeLessThan(2_000);
});
it("admits exactly 20 generated XSD levels and omits level 21", () => {
for (const [levels, noticeExpected] of [
[20, false],
[21, true],
] as const) {
const workspace = inspectWorkspace([
{ name: `depth-${levels}.xsd`, source: xsdChain(levels) },
]);
const sample = generateSample(workspace);
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(parsed.querySelectorAll("*")).toHaveLength(20);
expect(/level safety bound/iu.test(sample.notices.join("\n"))).toBe(
noticeExpected,
);
}
});
it("keeps truncated XSD astral literals well formed", () => {
const value = `a${"😀".repeat(600)}`;
const workspace = inspectWorkspace([
{
name: "unicode.xsd",
source: `<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"><xs:simpleType name="Unicode"><xs:restriction base="xs:string"><xs:enumeration value="${value}"/></xs:restriction></xs:simpleType><xs:element name="root" type="Unicode"/></xs:schema>`,
},
]);
const sample = generateSample(workspace);
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(hasUnpairedSurrogate(parsed.documentElement.textContent ?? "")).toBe(
false,
);
});
it("omits duplicate and fallback-colliding XSD attributes", () => {
const workspace = inspectWorkspace([
{
name: "attributes.xsd",
source: `<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"><xs:complexType name="WithAttributes"><xs:attribute name="same"/><xs:attribute name="same"/><xs:attribute name="1bad"/><xs:attribute name="2bad"/></xs:complexType><xs:element name="root" type="WithAttributes"/></xs:schema>`,
},
]);
const sample = generateSample(workspace);
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(parsed.documentElement.attributes).toHaveLength(2);
expect(parsed.documentElement.hasAttribute("same")).toBe(true);
expect(parsed.documentElement.hasAttribute("attribute")).toBe(true);
expect(sample.notices.join("\n")).toMatch(/duplicate.*attribute/iu);
});
it("caches a wide inline XSD type across repeated rendering", () => {
const annotations = Array.from(
{ length: 20_000 },
() => "<xs:annotation/>",
).join("");
const branches = Array.from(
{ length: 50 },
(_, index) => `<xs:element name="branch${index}" type="Branch"/>`,
).join("");
const workspace = inspectWorkspace([
{
name: "wide-inline.xsd",
source: `<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"><xs:complexType name="Branch"><xs:sequence><xs:element name="leaf">${annotations}<xs:simpleType><xs:restriction base="xs:string"/></xs:simpleType></xs:element></xs:sequence></xs:complexType><xs:complexType name="Root"><xs:sequence>${branches}</xs:sequence></xs:complexType><xs:element name="root" type="Root"/></xs:schema>`,
},
]);
const started = performance.now();
const sample = generateSample(workspace);
const elapsed = performance.now() - started;
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(elapsed).toBeLessThan(2_000);
});
it("inventories Relax NG and Schematron without executing expressions", () => {
const rng = inspectWorkspace([
{
@@ -287,6 +754,157 @@ describe("schema languages", () => {
).toBe(true);
});
it("retains a direct-root Relax NG element in the generated sample", () => {
const workspace = inspectWorkspace([
{
name: "direct.rng",
source: `<element xmlns="http://relaxng.org/ns/structure/1.0" name="root"><text/></element>`,
},
]);
const sample = generateSample(workspace);
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(parsed.documentElement.localName).toBe("root");
expect(parsed.documentElement.textContent).toBe("string");
});
it("never emits more than the exact Relax NG sample node budget", () => {
const repeatedElements = Array.from(
{ length: SCHEMA_LIMITS.sampleNodes },
() => `<element name="node"><text/></element>`,
).join("");
const workspace = inspectWorkspace([
{
name: "bounded.rng",
source: `<grammar xmlns="http://relaxng.org/ns/structure/1.0"><start><element name="root"><group>${repeatedElements}</group></element></start></grammar>`,
},
]);
const sample = generateSample(workspace);
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(parsed.querySelectorAll("*")).toHaveLength(
SCHEMA_LIMITS.sampleNodes,
);
});
it("applies the aggregate text budget through repeated Relax NG refs", () => {
const references = Array.from(
{ length: SCHEMA_LIMITS.sampleNodes },
() => `<ref name="large"/>`,
).join("");
const workspace = inspectWorkspace([
{
name: "text-budget.rng",
source: `<grammar xmlns="http://relaxng.org/ns/structure/1.0"><define name="large"><value>${"x".repeat(SCHEMA_LIMITS.sampleValueChars)}</value></define><start><element name="root"><group>${references}</group></element></start></grammar>`,
},
]);
const sample = generateSample(workspace);
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(parsed.documentElement.textContent?.length ?? 0).toBeLessThanOrEqual(
SCHEMA_LIMITS.sampleTextChars,
);
expect(sample.notices.join("\n")).toMatch(/text safety bound/iu);
});
it("bounds amplified Relax NG pattern visits independently of output", () => {
const width = 300;
const emptyPatterns = Array.from({ length: width }, () => "<empty/>").join(
"",
);
const references = Array.from(
{ length: width },
() => `<ref name="fanout"/>`,
).join("");
const workspace = inspectWorkspace([
{
name: "amplified-work.rng",
source: `<grammar xmlns="http://relaxng.org/ns/structure/1.0"><define name="fanout"><group>${emptyPatterns}</group></define><start><element name="root"><group>${references}</group></element></start></grammar>`,
},
]);
const started = performance.now();
const error = captureError(() => generateSample(workspace));
const elapsed = performance.now() - started;
expect(error.message).toMatch(/sample budget was exhausted/iu);
expect(elapsed).toBeLessThan(2_000);
});
it("caches wide reused Relax NG pattern containers", () => {
const ignored = Array.from({ length: 19_000 }, () => "<empty/>").join("");
const references = Array.from(
{ length: 1_000 },
() => '<ref name="wide"/>',
).join("");
const workspace = inspectWorkspace([
{
name: "wide-reused.rng",
source: `<grammar xmlns="http://relaxng.org/ns/structure/1.0"><define name="wide"><choice><empty/></choice>${ignored}</define><start><element name="root"><group>${references}</group></element></start></grammar>`,
},
]);
const started = performance.now();
const sample = generateSample(workspace);
const elapsed = performance.now() - started;
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(parsed.documentElement.localName).toBe("root");
expect(elapsed).toBeLessThan(2_000);
});
it("counts generated Relax NG levels rather than wrapper patterns", () => {
for (const [levels, noticeExpected] of [
[20, false],
[21, true],
] as const) {
const workspace = inspectWorkspace([
{ name: `depth-${levels}.rng`, source: relaxNgChain(levels) },
]);
const sample = generateSample(workspace);
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(parsed.querySelectorAll("*")).toHaveLength(20);
expect(/level safety bound/iu.test(sample.notices.join("\n"))).toBe(
noticeExpected,
);
}
});
it("keeps truncated Relax NG astral literals well formed", () => {
const value = `a${"😀".repeat(600)}`;
const workspace = inspectWorkspace([
{
name: "unicode.rng",
source: `<element xmlns="http://relaxng.org/ns/structure/1.0" name="root"><value>${value}</value></element>`,
},
]);
const sample = generateSample(workspace);
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(hasUnpairedSurrogate(parsed.documentElement.textContent ?? "")).toBe(
false,
);
});
it("omits duplicate and fallback-colliding Relax NG attributes", () => {
const workspace = inspectWorkspace([
{
name: "attributes.rng",
source: `<element xmlns="http://relaxng.org/ns/structure/1.0" name="root"><attribute name="same"/><attribute name="same"/><attribute name="1bad"/><attribute name="2bad"/><text/></element>`,
},
]);
const sample = generateSample(workspace);
const parsed = new DOMParser().parseFromString(sample.output, "text/xml");
expect(parsed.querySelector("parsererror")).toBeNull();
expect(parsed.documentElement.attributes).toHaveLength(2);
expect(parsed.documentElement.hasAttribute("same")).toBe(true);
expect(parsed.documentElement.hasAttribute("attribute")).toBe(true);
expect(sample.notices.join("\n")).toMatch(/duplicate.*attribute/iu);
});
it("parses OpenAPI YAML, inventories operations, and derives component samples", () => {
const workspace = inspectWorkspace([
{
@@ -320,6 +938,38 @@ components:
name: "string",
});
});
it("refuses an impossible OpenAPI component sample", () => {
const workspace = inspectWorkspace([
{
name: "impossible-openapi.json",
source: JSON.stringify({
openapi: "3.1.0",
info: { title: "Impossible", version: "1" },
paths: {},
components: { schemas: { Impossible: false } },
}),
},
]);
expect(() => generateSample(workspace)).toThrow(/no valid sample/iu);
});
it("applies the same 20-level depth ceiling to OpenAPI samples", () => {
const workspace = inspectWorkspace([
{
name: "deep-openapi.json",
source: JSON.stringify({
openapi: "3.1.0",
info: { title: "Deep", version: "1" },
paths: {},
components: { schemas: { Deep: nestedJsonSchema(21) } },
}),
},
]);
const sample = generateSample(workspace);
expect(jsonValueDepth(JSON.parse(sample.output))).toBe(20);
expect(sample.notices.join("\n")).toMatch(/level safety bound/iu);
});
});
describe("conservative comparison", () => {