Language Bindings
Every interface wraps the same Store. Understanding how each binding bridges to the core explains
why the APIs look the way they do, how errors propagate, and what each layer owns.
flowchart TB
PYAPP["Python code"] --> PYO3["PyO3 classes<br/>(infrastore_py)"]
JLAPP["Julia code"] --> JLPKG["InfraStore.jl"]
JLPKG -->|"ccall"| CABI["C ABI<br/>(infrastore_ffi)"]
RUSTAPP["Rust client code"] --> RC["RemoteClient"]
RC -->|"gRPC / HTTP2"| GS["gRPC server"]
PYO3 --> STORE["Store"]
CABI --> STORE
GS --> STORE
style STORE fill:#28a745,color:#fff
style PYO3 fill:#17a2b8,color:#fff
style CABI fill:#9558b2,color:#fff
style JLPKG fill:#9558b2,color:#fff
style GS fill:#ffc107,color:#000
style RC fill:#ffc107,color:#000
Python (PyO3)
infrastore-py uses PyO3 to expose Store as native Python classes in a module
importable as infrastore. The binding:
- Converts Python
datetime/timedeltatochronotypes and NumPy arrays (any shape) toTypedArrays at the boundary, supporting the full dtype set (f64,f32, the integer widths,bool). - Translates the typed
TimeSeriesErrorvariants into a Python exception hierarchy rooted atTimeSeriesError(NotFoundError,DuplicateTimeSeriesError,InvalidParameterError,IntegrityError,ReadOnlyStoreError). - Builds an
abi3-py311wheel, so one wheel works across CPython 3.11+ without recompiling. - Converts a static series to a
pyarrow.Tablewithto_arrow(), behind the optionalarrowextra. This is the one place a binding reaches past numpy, and it is optional for that reason: pyarrow is several times the size of the wheel that would pull it in. What makes Arrow worth the seam is that itstimestamp(unit, tz)is the same shape as the store's own model — an instant plus the spelling it was written in — so a table keeps a distinction pandas would flatten.
The metadata side is owned entirely by Rust; Python never touches SQLite directly. See the Python guide and Python API reference.
Julia (C ABI)
Julia does not call Rust directly. Instead, infrastore-ffi compiles a C-compatible cdylib with an
opaque-handle API, and InfraStore.jl ccalls into it.
flowchart LR
JL["InfraStore.jl<br/>structs hold Ptr{Cvoid}"] -->|"ccall infrastore_store_*"| LIB["libinfrastore_ffi"]
LIB --> STORE["Store"]
LIB -.->|"infrastore_last_error_message"| JL
style JL fill:#9558b2,color:#fff
style LIB fill:#6f42c1,color:#fff
style STORE fill:#28a745,color:#fff
The conventions that shape the Julia API:
- Opaque handles.
InfraStoreandInfraStoreKeyare pointers; the Julia structs wrap them and register finalizers (close!,_finalize_key) that call the matchingts_*_freefunction. - Status codes plus thread-local error messages. Every C function returns an
int32_tcode. On a non-zero code, Julia callsinfrastore_last_error_messageto retrieve the detail string and raises the matching Julia exception type. - Out-parameters and caller-owned buffers. Arrays come back through an out-pointer plus a length
and a dtype code; Julia copies them into a
Vector{T}for the requested element type and frees the Rust buffer with the deallocator matching the buffer's element type —infrastore_buffer_free_f64,infrastore_buffer_free_u8,infrastore_buffer_free_i64, orinfrastore_buffer_free_u64(shape/dims buffers). - Features cross as JSON. Julia serializes the feature dict to a JSON string, which the FFI
layer parses into a
Featuresmap. - Forecasts are wrapped.
InfraStore.jlexposesDeterministic/Probabilistic/Scenariosstructs passed to the genericadd_time_series!, id-addressedread_by_idgetters, andtransform_single_time_series!, so all four forecast types are usable from Julia. - Bulk reads use a result handle.
read_by_idsreads many fullSingleTimeSeriesat once: the FFI fetches them in one decompress-once pass per dataset into aInfraStoreBulkReadHandle(infrastore_store_bulk_read_single), and Julia reads each element out, then frees the handle. Python'sstore.read_by_idsexposes the same operation directly.read_by_idsaddresses the same read by catalog association id and fills the same handle, so both reads decode by the same route:infrastore_bulk_result_item_namehands each item's name back beside its values, asinfrastore_bulk_result_item_typedoes its type. Managed bulk writes already take the fast block-write path through the existing batch /add_time_series_bulkAPIs.
InfraStore.jl loads the cdylib from the INFRASTORE_LIB environment variable when it is set, and
otherwise from the libinfrastore_ffi artifact its Artifacts.toml pins to the matching GitHub
Release (see Integrate with Julia). See the
Julia guide, the C ABI reference, and the
Julia API reference.
InfrastructureSystems.jl Integration
The model was shaped to drop into InfrastructureSystems.jl: owners are identified by integer
component identifiers (i64), owner categories map to Component / SupplementalAttribute, and
features accept string values so InfrastructureSystems.jl's feature dictionaries round-trip
unchanged. The FFI exposes an attribute-based existence probe (infrastore_store_has_any_by_filter)
and removal (infrastore_store_remove_by_ids), a whole-record metadata read
(infrastore_store_get_metadata_by_id, reachable from attributes through
infrastore_store_list_metadata), and a hash-based array fetch
(infrastore_store_get_array_by_hash) so an InfrastructureSystems.jl-side store can keep its own
key objects and reach the array layer directly.
gRPC Server and Client
infrastore-server wraps a Store in a tonic gRPC service generated from infrastore-proto. It
exposes a read-only slice of the API and adds optional API-key auth. The matching async
RemoteClient mirrors the read methods and maps gRPC Status codes back to
TimeSeriesError::ConnectionError, so remote calls surface the same error type as local ones.
Writes are deliberately not exposed over gRPC — they require local filesystem access. The server is for fan-out reads of an existing store. See the gRPC Server guide and the gRPC API reference.
CLI (infrastore)
infrastore-cli builds the infrastore binary, a thin wrapper over the core Store for use from a
terminal. Unlike the gRPC server it is not read-only: it opens the on-disk .h5 + .h5.sqlite
pair directly and supports both reads and writes. Its shape:
- CSV in, store out. Numeric values come from a CSV; the metadata that does not fit a flat grid (owner, name, type, dtype, resolution, timestamps, units, features) is described in a descriptor JSON. All six dtypes and all six writable types are supported, forecasts included.
- A global
-f/--formatselectstable(default),json,jsonl, orcsv. Read commands render their results in it; write commands report their outcome in it (prose undertable, a one-object status document underjson/jsonl). Onlytemplateignores it. - Store access is isolated. All store opening lives behind one module, so a future remote/gRPC mode can be added without touching the command handlers; today there is no remote mode.
See the CLI guide and the CLI reference.
What Every Binding Shares
| Concern | Single source of truth |
|---|---|
| Types & validation | infrastore-core (Store, TimeSeriesId, Features) |
| On-disk format | Hdf5Backend + MetadataStore — identical regardless of caller |
| Hashing | array_hash / features_hash — the cross-language contract |
| Error taxonomy | TimeSeriesError, re-projected into each language's idiom |
A file written by Python reads identically from Julia, Rust, or the server, because none of the bindings reimplement storage — they all funnel through the one core.
Feature Coverage Varies by Binding
The bindings funnel through one core, and the surface is now broadly consistent. Both static series types are available everywhere (read+write, except the read-only gRPC server), and forecasts read back across every interface. The remaining asymmetry is that the read-only gRPC server does not accept any writes:
| Capability | Rust core | C ABI | Python | Julia | CLI | gRPC |
|---|---|---|---|---|---|---|
SingleTimeSeries r/w | ✅ | ✅ | ✅ | ✅ | ✅ | read-only |
NonSequentialTimeSeries r/w | ✅ | ✅ | ✅ | ✅ | ✅ | read-only |
PersistentTimeSeries r/w | ✅ | ✅ | ✅ | ✅ | ✅ | read-only |
dtypes beyond f64 | ✅ | ✅ | ✅ | ✅ | ✅ | read-only |
| Create forecasts | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ |
| Read forecast values | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Forecast metadata / counts | ✅ | ✅ | ✅ | ✅ | ✅ | list/counts |
| Readers (columnar sweep) | ✅ | ✅ | ✅ | ✅ | grid | ❌ |
| Association catalogs | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ |
| Store attributes | ✅ | ✅ | ✅ | ✅ | store-attr | read-only |
| Materialized timestamps | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ |
from_timestamps (verified) | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ |
Arrow tables (to_arrow) | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
| Parquet files | crate | ❌ | ❌ | ❌ | -f parquet | ❌ |
Store summary (show) | ❌ | ❌ | ✅ | ❌ | store-info | ❌ |
| Forecast windows as Arrow | ❌ | ❌ | Deterministic | ❌ | ❌ | ❌ |
The only gap is by design: writes (including forecasts added through add_time_series) require
local filesystem access, so the read-only gRPC server serves forecast reads but not writes.
show() is Python-only for now: it is a REPL affordance, and the REPL each binding is used from
already has one of its own — Julia has Base.show, and the CLI has store-info plus the list
family. It composes existing catalog aggregate queries and adds no core API, so any binding that
wants it can grow one without a change underneath.
Parquet lives in a crate of its own, infrastore-parquet, which the CLI depends on behind a
cargo feature that is on by default — the infrastore binary anyone installs can read and write
Parquet, because handing an analyst a file for DuckDB or polars is an ordinary reason to reach for
the CLI. The feature stays switchable (--no-default-features --features vendored builds a lean
binary), and the line it draws is between the binary and the libraries: infrastore-core,
infrastore-py, and infrastore-ffi never link Arrow, which cargo tree --edges normal on each is
the check for. The CLI is the only surface that reads and writes Parquet files
(export -f parquet, add --parquet): a normalized, partitioned layout of two files per
partition — a values file holding each distinct array once and a series file holding the catalog
rows that name it, joined on (data_hash, time_axis) — specified in the
Parquet layout reference. Python's to_arrow() / from_arrow are
a different, in-memory thing — one two-column table per series, with the descriptors in the schema
metadata rather than in columns — and are not a reader or writer for the CLI's files; a Python user
who wants one writes the per-series table with pyarrow.parquet, or hands the CLI's directory to
DuckDB or polars. The relationship between the two is one sentence: a series file's columns are
to_arrow()'s metadata keys turned into columns, and the values file is its two columns keyed by
the array. Nothing else has it: the C ABI and Julia would need the whole Arrow tree in the cdylib
for a format their host languages already have readers for, and the gRPC server serves values, not
files.
Materialized timestamps and from_timestamps both run in the core and reach Julia through
two stateless ABI entry points, infrastore_grid_timestamps and infrastore_infer_period. That
matters more than it looks: Julia is the one binding whose date library has calendar arithmetic of
its own, and whose TimeZones overload steps a local clock the core deliberately does not — so a
binding-side reimplementation would agree with the core only by luck. There is one implementation of
"which instants does this series contain" in the project, and it is Period::add_to. Arrow
tables are Python-only because Arrow is where the Python data ecosystem meets; the Julia
counterpart would be a Tables.jl interface, which is a different contract and not yet asked for. A
Deterministic converts through to_arrow_windows() into one table per window rather than one
table, because its two grids — windows stepping by interval, rows stepping by resolution —
overlap; Probabilistic and Scenarios wait on a decision about how to spell their third axis.