Qortex Tools
Developer and Format/Multi-sample/4 drafts/Runs in your browser

JSON Schema Generator

Paste one or more JSON payloads and infer a JSON Schema. Detects string formats, warns about duplicate keys and BigInt precision. Nothing uploads.

Why this matters in real testing work

JSON Schema Generator: Infer Schemas From Sample Payloads

Most JSON Schema generators ask you to describe your data by hand. This tool works the other way around: paste real JSON payloads and let the inference engine figure out the schema for you. Everything runs in your browser. Nothing is uploaded, no account is needed, and your data stays on your machine.

How to use it

Using the tool takes five steps.

  1. Paste a JSON sample. Drop a raw JSON object or array into the input area. The tool accepts any valid JSON, from a two-field object to a deeply nested API response.

  2. Add more samples (optional). Click "Add sample and paste another" to include additional payloads. Each sample contributes to the inferred schema. The more samples you provide, the more accurate the output becomes, because the engine can distinguish required fields from optional ones and narrow down type unions.

  3. Pick a draft. The dropdown offers four JSON Schema specifications: Draft 04, Draft 07, Draft 2019-09, and Draft 2020-12 (the default). Each draft produces the correct $schema URI so validators know which specification to apply.

  4. Toggle format detection. When the "Detect string formats" checkbox is checked, the engine tests every string value against six common patterns: date-time, email, URI, UUID, IPv4, and IPv6. A format annotation appears on a property only when every value of that property across every sample matches the same pattern.

  5. Click Generate. The engine runs and produces a JSON Schema in the output panel. From there you can copy the schema to your clipboard, download it as a .json file, or continue refining by adding more samples and regenerating.

Multi-sample inference

The inference engine uses a Map-Reduce architecture. In the Map phase, each sample is parsed independently into a draft-agnostic internal representation. In the Reduce phase, the representations are merged pair by pair until a single schema remains.

The merge follows three rules:

  • Required field intersection. A property is marked required only when it appears in every sample. A property missing from even one sample becomes optional.
  • Type widening. When the same property holds different types across samples, the schema lists all observed types as a union. If one sample has an integer and another has a float, the engine widens to number rather than listing both.
  • Format consensus. A format annotation survives the merge only when every string value of that property matches the same pattern. One disagreeing value suppresses the annotation entirely.

This means you can paste a handful of real API responses and get a schema that reflects the shape your consumers actually see, including optional fields and mixed types that a single sample would miss.

Draft versions

The tool supports four JSON Schema specifications:

  • Draft 04 (http://json-schema.org/draft-04/schema#): the oldest supported draft. Some legacy validators and codegen tools still require it.
  • Draft 07 (http://json-schema.org/draft-07/schema#): widely adopted and the baseline for many contract testing tools.
  • Draft 2019-09 (https://json-schema.org/draft/2019-09/schema): introduced vocabulary and improved $ref handling.
  • Draft 2020-12 (https://json-schema.org/draft/2020-12/schema): the current stable specification and the tool's default.

In version one, the generated keywords (type, properties, required, items, format) are identical across drafts. The difference is the $schema URI at the root, which tells validators which specification to apply. The serializer is structured to support draft-specific keywords in future versions without architectural changes.

String format detection

When format detection is enabled, the engine tests every string value against six regular expression patterns:

  • date-time: ISO 8601 timestamps like 2024-01-15T10:30:00Z or 2024-03-22T14:15:00+05:30.
  • email: addresses matching the common user@domain.tld pattern.
  • uri: HTTP and HTTPS URLs.
  • uuid: RFC 4122 UUIDs with version digits 1 through 5.
  • ipv4: dotted-decimal addresses with valid octet ranges.
  • ipv6: full-form colon-separated addresses.

The consensus rule ensures accuracy: a format annotation is added only when every occurrence of a property across every sample matches the same pattern. If you paste ten responses and one of them has a plain string where the others have an email, the engine omits "format": "email" for that property. You can disable format detection entirely with the checkbox if you prefer a schema without format annotations.

Data quality warnings

The tool catches two problems that most schema generators ignore:

Duplicate keys. The JSON specification allows an object to contain the same key more than once. Most parsers silently keep the last occurrence and drop the rest, which means you never notice. This tool scans the raw text before inference and flags every duplicate key with a JSON Pointer, the affected sample number, and a human-readable explanation.

Integers past 2 to the 53rd. JavaScript's Number type loses precision above Number.MAX_SAFE_INTEGER (9007199254740991). An integer like 9007199254740993 is silently rounded during parsing, and the rounded value is what ends up in the schema. The tool detects these values in the raw text and warns you before inference runs.

Both warnings appear alongside the generated schema. The schema itself is still produced, because the warnings are informational: they tell you about data quality issues in your input, not errors in the generated schema.

What the tool does not generate

Version one focuses on the core inference features. The following are deliberately omitted:

  • enum keyword. The engine does not collect observed values and emit them as enumerations. Enums are better specified by hand, based on domain knowledge the engine cannot infer.
  • additionalProperties keyword. The engine does not restrict objects to their observed properties. Closing schemas should be a conscious design decision, not an inference artifact.
  • $ref and definitions. Sub-schemas are always inlined. Deduplication into $defs or definitions is planned for a future version.
  • Tuple inference. All arrays are treated as homogeneous lists with a single items schema. The engine does not emit prefixItems or position-dependent tuple types.
  • File upload. Input is by paste only. There is no file picker, drag-and-drop, or upload endpoint. This keeps the privacy guarantee simple: nothing leaves the browser because there is no channel for it to leave through.

Privacy

Schema inference runs entirely in a web worker in your browser. The tool makes no network requests while generating. Your JSON payloads never leave your machine, and no telemetry captures input or output content. The worker can be cancelled at any time using the Cancel button.

QL
Qortex Lab
Tools and guides for software testing
Common questions

Frequently asked questions

Answers to what people ask most. These are the same questions and answers used in the page's structured data.

What JSON Schema drafts are supported?+

Draft 04, Draft 07, Draft 2019-09 and Draft 2020-12. The default is 2020-12, the latest stable specification. Each draft produces the correct $schema URI.

Does my data leave the browser?+

No. Schema inference runs entirely in a web worker in your browser. Nothing is uploaded, no account is needed, and the tool makes no network requests while generating.

How does multi-sample inference work?+

Paste any number of JSON samples and the tool merges them into one schema. Properties present in every sample are marked required. Conflicting types become union types. A string format is emitted only when every value of that property across every sample matches the same format pattern.

What do the duplicate key and BigInt warnings mean?+

JSON allows an object to carry the same key twice, and most parsers silently keep the last occurrence. The tool detects this before inference. Similarly, integers above 2 to the 53rd lose precision during JavaScript parsing. Both are reported as warnings alongside the generated schema.

Related tools