Skip to content

Remove malformed records

Real JSON sometimes holds records that are broken in a recognisable way: an entry that carries an error field instead of its data, or a bare string where every other row has an object. genson normalises these like any other record, so they reach the output, and the fields only they have (error, a promoted …__string) widen the schema for every row.

prune on normalise_from_parquet names those fields. Every record holding one is removed during normalisation and written to a separate file, and the fields are left out of the output schema.

Example: sensor readings with failures

import polars as pl

pl.DataFrame({
    "sensor": ["a", "b", "c"],
    "data": [
        '{"readings": [{"t": 1, "value": 3.2, "calibration": {"offset": 0.1}}, {"t": 2, "error": "timeout"}]}',
        '{"readings": [{"t": 1, "value": 2.8, "calibration": "pending"}]}',
        '{"readings": [{"t": 1, "value": 1.5}]}',
    ],
}).write_parquet("sensors.parquet")

Sensor a has a failed reading holding error. Sensor b's calibration is the string "pending" where the others have an object, so genson promotes it to the field calibration__string. Prune both:

from polars_genson import normalise_from_parquet

normalise_from_parquet(
    "sensors.parquet",
    column="data",
    output_path="clean.parquet",
    typed=True,
    keep_columns=["sensor"],
    prune={"error", "calibration__string"},
    prune_output_path="pruned.parquet",
)
print(pl.read_parquet("clean.parquet").unnest("data"))
print(pl.read_parquet("pruned.parquet"))
shape: (3, 2)
┌────────┬─────────────────┐
│ sensor ┆ readings        │
│ ---    ┆ ---             │
│ str    ┆ list[struct[3]] │
╞════════╪═════════════════╡
│ a      ┆ [{1,3.2,{0.1}}] │
│ b      ┆ null            │
│ c      ┆ [{1,1.5,null}]  │
└────────┴─────────────────┘
shape: (2, 3)
┌────────┬────────────────┬─────────────────────────────────┐
│ sensor ┆ path           ┆ value                           │
│ ---    ┆ ---            ┆ ---                             │
│ str    ┆ str            ┆ str                             │
╞════════╪════════════════╪═════════════════════════════════╡
│ a      ┆ ["readings",1] ┆ {"t":2,"error":"timeout"}       │
│ b      ┆ ["readings",0] ┆ {"t":1,"value":2.8,"calibratio… │
└────────┴────────────────┴─────────────────────────────────┘

What gets removed

  • A record is pruned if it holds a non-null value for a named field, at any depth.
  • Names are the fields of the inferred schema, so a scalar promoted to a record goes by its promoted name, as it appears in the schema (calibration__string).
  • A pruned record that is a field of another record takes that record with it, since the record is incomplete without it. This goes on up to an array element or map entry, which is removed on its own. Above, the pending calibration takes its reading with it.
  • An array or map left empty by pruning is removed from its array or map, or is null as a record field. Sensor b's readings are null for this reason, while the row itself is kept. An array or map already empty in the input is normalised as usual.
  • The named fields are left out of the output schema.

The pruned values

Nothing is dropped silently. prune_output_path gets one row per pruned value:

  • the keep_columns of its row, to tell which row it came from;
  • path: the keys and array indices from the row's root to the value, as a JSON array;
  • value: the value as it was in the input, as JSON.

prune and prune_output_path must be given together.

Finding the names

Infer the schema first and look for fields that only the malformed records have: an error or message, or a …__string beside a record's usual fields. See Mixed types for how promoted names are formed.