How inference works¶
genson infers a schema in two passes. The first reads the data and builds one JSON Schema that every row fits. The second rewrites that schema: it decides which objects are maps, merges the different shapes a value takes, and puts union types in a fixed order. This page describes both. The pages before it describe the outcome; this one describes the algorithm, for when you need to know why a schema came out as it did, or what a change to the data will do to it.
Pass 1: build a schema from the data¶
One tree per position in the data¶
genson parses each JSON string with simd-json and records what it sees in a tree with one node per position in the data: the root, each field of an object, the items of an array, the values of a map. A node records each kind of value seen at its position:
| Seen at the position | Recorded |
|---|---|
| an object | its fields, in the order first seen; the keys present in every object |
| an array | one node for all of its items |
| a number | integer, until a number that is not an integer is seen: then number |
| a string, boolean, null | the type |
A position that held more than one kind of value gets a schema for each: scalars become a
list of types ({"type": ["null", "string"]}), and an object or array alongside others
becomes an anyOf. A field is required when every object seen at that position held
it.
wrap_root writes each document inside an object of one field ({"<field>": <document>})
before it is parsed.
Many strings: one tree each, then merged¶
With 10 or more strings, each string gets a tree of its own on the thread pool. The trees
are then merged, in input order, into one: fields are united in the order first seen,
the required keys are those required in every tree, and a number position becomes
number if it is number in any tree. Merging two trees gives the same schema as
reading all of their data into one tree.
- Repeated schemas are merged once. Each string's tree is hashed as it is built, and a
tree whose schema equals one already merged is skipped. Rows of the same shape are
common, so this saves most of the merging. (The
processed_countthat inference reports is the number of distinct schemas merged.) - The merge runs in parallel across fields. When the trees being merged are all objects, each field of the root is merged on its own thread, and so on down the tree while there are enough trees to make it worthwhile. Each field still takes its values in input order, so the result does not depend on the number of threads.
max_buildersbounds memory. Strings are processed in chunks ofmax_builders; each chunk's trees are merged before the next chunk is parsed (see Process large inputs).- Fields forced to be maps (
force_field_types={"field": "map"}) are turned into maps in each string's tree before the merge, so that their values are merged across strings as map values rather than as fields.
With fewer than 10 strings, each string is read into the same tree in turn. Within one string, the documents of NDJSON (and arrays of 10 or more items anywhere) are read on several threads and merged the same way.
Pass 2: rewrite the schema¶
The merged tree is written out as a JSON Schema, and genson walks it from the root, carrying each field's name down so that options naming fields apply. At each object, the first of these rules that applies decides what happens:
- A field in
force_scalar_promotionholding a scalar becomes a record with one field,<field>__<type>(for exampledatavalue__string). - A field in
force_field_typesbecomes a map ("map"), whose value schema is the values' common schema, unified as below if they differ, or elsestring; or stays a record ("record"). - An
anyOfof scalars and records, withwrap_scalarson (the default), becomes one record: each scalar is promoted to a record of one field<field>__<type>, and the records are unified. - An object that was empty in every row, with
map_threshold=0, becomes a map of null (a record with no fields cannot be written to Parquet). - An object holding a field named in
force_parent_field_typesas"record", or a field inforce_scalar_promotion, stays a record. -
An object becomes a map when it has at least
map_thresholdkeys, is not the root (unlessno_root_map=False), and- all of its values are the same record; or else
- it has at most
map_max_required_keysrequired keys (if that is set), and its values' schemas are the same once nullability is set aside or, withunify_maps, they unify.
The map's value schema is that common schema.
Then genson goes on into the object's fields, an array's items and a map's values.
Unifying value schemas¶
Unification merges the schemas of a candidate map's values into one, when they are of one kind:
| Values | Unified as |
|---|---|
| arrays | an array of the items unified; an array only ever seen empty fits any array |
| maps (and empty records) | a map of the values unified |
| records | a record of every field seen: unified field by field, required and non-null where every record has the field, nullable otherwise |
| scalars of one type | that type, nullable |
Within a record field, two schemas that differ only in being nullable merge into the
nullable one, and with wrap_scalars a scalar that meets a record (or a scalar of
another type) is promoted under <field>__<type>. Array items unify the same way, and
there a scalar among records is promoted too. Values of different kinds that none of
this reconciles (an array and a record, say) do not unify, and the object stays a
record. Nor is an object unified when one of its keys is named in no_unify.
Union order¶
Finally, every list of types is put in a fixed order: null first, then containers
(map, array, object), then scalars from the narrowest (boolean, integer, number,
string). A ["null", T] pair keeps its order.