Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Element Types

Every stored array carries an element_type: what its elements mean, and — for the composite kinds — how one timestep's values are laid out across the array's trailing dimensions.

It is a first-class, store-owned concept, not a binding convention. A Julia PiecewiseLinearData and a Python list of {"x": …, "y": …} dicts are both piecewise_linear here, so a consumer written in either language can decode an array without knowing which language wrote it.

element_type replaces a separate physical dtype column: the dtype of the stored bytes is derived from it. TypedArray still carries a dtype, because it describes bytes; the element type lives on the association metadata and on the write API, where interpretation belongs.

Canonical string form

The element type travels as a string: the element_type column in the SQLite catalog, a UTF-8 string across the C ABI, and a string field over gRPC. A parameterized grammar does not fit an integer code, so unlike dtype there is no numeric encoding.

f64 | f32 | i64 | i32 | i16 | i8 | u64 | u32 | u16 | u8 | bool
tuple(N,dtype)          e.g. tuple(3,f64)
linear_function
quadratic_function
piecewise_linear
piecewise_step

The names are deliberately language-neutral: piecewise_linear, not PiecewiseLinearData; tuple(3,f64), not NTuple{3, Float64}. Each binding maps them to its own domain types.

Encoding

The first array dimension is time (the per-window horizon for forecasts). The trailing dimensions are the per-step element shape, determined by the element type:

element_typeelement shaperow layout (one timestep)
a dtype spelling[]the value
tuple(N,T)[N]t1 … tN
linear_function[2]proportional, constant
quadratic_function[3]quadratic, proportional, constant
piecewise_linear[1 + 2*w]n, x1, y1, …, xn, yn, zero-padded to width
piecewise_step[max(1, 2*w)]n, x1 … xn, y1 … y(n-1), zero-padded to width

w is the maximum point / x-coordinate count across the series (or across every window of a forecast). Ragged rows are self-describing through the leading count n, so decoding one row needs no global state.

A scalar element type still allows a dense per-step array — a Probabilistic forecast's percentile columns, say. The dense-array case and the tuple(N,T) case differ in meaning, not bytes: the tuple says the N values are one composite value, not N independent samples.

Forecasts stack windows in front of the per-step element shape: [H, count, *E] for a Deterministic, [P, H, count, *E] for a Probabilistic or Scenarios. The per-row scheme is unchanged; only how many leading axes precede it differs (TimeSeriesType::leading_dims).

Every function-data kind is stored as f64.

Validation

Because the store owns the element type, it can reject a write that contradicts it rather than storing the inconsistency blindly:

  • the array's dtype must equal the element type's physical dtype;
  • fixed-width kinds must have exactly their per-step dims ([2], [3], [N]);
  • ragged kinds must have exactly one trailing dim, of a width their layout can produce (odd for piecewise_linear; 1 or even for piecewise_step);
  • every ragged row's leading count n must be a non-negative whole number that fits the row.

Scalars are unconstrained in their trailing dims, since a dense per-step array is legitimate.

Every series carries a concrete element type — a constructor resolves it to plain scalars of the array's dtype, and declaring one replaces it — so the checks above run on every write, not only on writes that declared something. There is no "undeclared" state to fall back from.

The catalog is the source of truth

Nothing about element typing is recoverable from the HDF5 file. Storage records how many bytes an element occupies; element_type records what it means, and even the physical dtype is not fully recoverable (bool and u8 are the same byte). Every read therefore resolves the element type from the catalog first and tells the storage backend what to decode — the backend never infers it.

infrastore verify follows from this: it walks the arrays the catalog references, so a catalog row pointing at an array the file does not hold, or a row too malformed to name one, is reported. An array in the file that no association references is not checked — it is unreachable, and nothing records what its bytes mean.

Codecs

Each binding ships a reference codec between the stored bytes and per-timestep values:

  • Rustinfrastore_core::{decode, encode} over TypedArray + ElementType. Prefer the paired forms: every value type has a from_values constructor that encodes the values and declares the element type they imply, and TimeSeriesData::decoded_values reads them back. An element_type and the array it describes are two things a caller can get out of step — the store rejects the mismatch on write, but deriving both from one set of values means there is none to reject. encode_as is the declared-type encoder for the one series from_values cannot name: a tuple with no rows, whose arity lives in rows it does not have. See Element values.

  • Python — the paired forms first, as in Rust: every series type has a from_values classmethod that encodes the values and declares the element type they imply, and .decoded_values() on a series reads them back. Underneath sit infrastore.decode_element_values(array, element_type, leading_dims) and encode_element_values(values, element_type, leading_dims), for the cases the pair cannot name — an empty tuple(N,f64) series, whose arity lives in rows it does not have — and for decoding an array that arrived without a series around it.

    A Python payload carries no type tag of its own, unlike a Rust DecodedValues or a Julia Vector{PiecewiseLinear}, so from_values reads the element type off the shape of a row. The five shapes are disjoint, which makes that a decision rather than a guess:

    values entryelement type
    {"proportional": …, "constant": …}linear_function
    {"quadratic": …, "proportional": …, "constant": …}quadratic_function
    list[{"x": …, "y": …}]piecewise_linear
    {"x": list, "y": list}piecewise_step
    list[float] of length Ntuple(N,f64)

    element_type= is still accepted, as an assertion rather than an override: it raises if it disagrees with the values. Where the values name nothing it is the only thing to go on — an empty values, or rows that are all empty and read equally as a pointless curve or a zero-arity tuple — and without it those are refused, naming the remedy. The one series no declaration reaches is an empty tuple(N,f64), whose arity lives in rows it does not have.

  • JuliaInfraStore.encode_element_values / decode_element_values, over the value types LinearFunction, QuadraticFunction, PiecewiseLinear and PiecewiseStep. Decode takes a types keyword, so a consumer with its own domain types — InfrastructureSystems.jl's FunctionData — decodes straight into them and pays no conversion; encode is three small generic functions it extends instead. The names follow the wire vocabulary (PiecewiseLinear, not PiecewiseLinearData) so that using InfraStore, InfrastructureSystems is not an ambiguity error.

The CLI decodes composite rows for get -f json, under an element_values key alongside the raw values. Its CSV output stays packed on purpose: that form is what add reads back, so it has to stay the store's own layout rather than a rendering of it.

Extending the codec

A consumer with its own domain types does not have to convert at the boundary. In Julia the two directions extend differently, because they start from different things:

  • Decoding starts from an element_type string, so the type to build is chosen by name: decode_element_values(...; types = ...), and read_by_id(store, id; types = ...).
  • Encoding starts from a value, so it is open dispatch: add element_type_tag, element_row_width and write_element_row! methods for your type and it packs directly.

is_element_values is the predicate the write path uses to tell "domain values to pack" from "numbers to store as they are", and it answers by asking whether those three methods exist — so opting a type in is exactly defining them, with nothing to register.

conformance/element_type_vectors.json at the repo root pins encoded bytes against expected decoded values for every element type, static and forecast. It is generated by infrastore-core's tests/element_type_conformance.rs and read by the Python and Julia codec tests, so every implementation is held to one definition of the encodings rather than to each other. Regenerate it with:

UPDATE_CONFORMANCE_VECTORS=1 cargo test -p infrastore-core --test element_type_conformance

A binding may reject what it cannot represent — the grammar allows tuple(4,i32), which the Julia binding does not map — but the store accepts the full grammar. A binding's codec is the other way round: it has to represent everything the store accepts, which is why the Julia value types take the zero- and one-point piecewise curves that a domain type like InfrastructureSystems.jl's PiecewiseLinearData rejects. A curve too short to interpolate is still a row a read has to hand back.

Because consumers go through the codecs, a future storage optimization (a true ragged layout with an offsets array instead of zero padding) can land behind this boundary without touching them.