Data binding¶
The YAML declares shapes; sources supplies the numbers, keyed by the names
the file declared:
import lpspec as lps
result = lps.solve(
'dispatch.yaml',
{'load': 'load.parquet', 'cost': cost_frame, 'p_max': p_max_frame},
)
What a parameter accepts¶
For a parameter declared dims: [d1, d2]:
- a parquet path;
- any table exposing the Arrow PyCapsule protocol — polars, pandas,
pyarrow, duckdb — with columns
d1, d2, value; - an
intorfloat, standing for every coordinate the parameter covers; - a
dictof label to value, for a parameter over one dimension; - a sequence — list, tuple,
np.ndarray— for a parameter over one dimension, positional against that dimension's index.
The last three are for models written out in Python. Each is dense, so each is
materialised at bind: one number over (snapshot, generator) becomes a row per
pair, and a value that really is constant is better declared dims: []. A
sequence is positional, so the dimension's labels have to come from somewhere
other than this parameter — one of the three sources below, which is what fixes
the order it is positional against.
pd.Series keeps its dims in an index rather than in columns, so it is
unwrapped first — but only if pandas is already imported, never by importing
it. An unnamed index binds positionally to the declared dims; a named one
binds by name in any order, and a name outside the declared dims raises rather
than being overwritten.
Tables in, arrays out. An xr.DataArray is a dense n-dimensional array
rather than a table, and neither lane reads one: pass array.to_series(),
whose index binds by name on both. Result.to_dataarray() is the way back out.
Everything on this list is read by the linopy lane
too, so one sources mapping goes to either.
Nothing on this path imports pandas, xarray or linopy on your behalf.
Where coordinates come from¶
Master coordinates are resolved per dimension before any parameter loads, from exactly one of:
- a key in
sources— a table carrying a column of that name, or a parquet path. The first occurrence of each value is its position; dimensions.<d>.valuesin the YAML — the labels written out in the file.
The two are exclusive, not ranked. A dimension the file declares and the caller also supplies is refused at bind, naming the declaration and the key that collided with it. The file owns that dimension's labels or the caller does, never both — so there is no precedence to remember, and no way for the file a reviewer reads to describe a model the caller quietly replaced. A model whose label set varies from run to run should not declare one.
A declared map is not on
that list either. lookups.<x>.values says how labels map, never which ones exist: a
map is a partial relation over the dimension, free to omit members and written
in whatever key order someone typed, and neither may decide an extent nor an
order that shift reads positionally. A map is instead
read against whichever of the two supplied the labels. Each map has one home
in the same way the labels do — the file, or a column of the caller's index —
and claiming both is the same refusal. Which leaves exactly one index with two
authors, one fact each: labels from the caller, maps from the file.
Reading a map against labels is not symmetric, because the two directions mean different things. A label no map mentions gets a null — the partial case, and what a relation over a dimension is entitled to be. A key matching no label is a typo, and refused: dropping it would place its terms nowhere while the model built and solved. Where the file declares the labels too, that same refusal happens at load with no data at all.
There is no third step. A dimension neither of the two supplies raises, and
labels are never read out of the parameters: they would be the definition,
so a mistyped label could not be told from a new one, and the index is also
what fixes label order, which shift reads
positionally.
The data contract¶
Both lanes bind by these rules, and tests/test_data_parity.py is what holds
them to it: the same malformed source, checked for the same verdict and — where
one defect has one repair — the same sentence.
A coordinate has a value, or it has no row. A row whose value is null or
NaN says both at once, and is refused at bind naming the parameter and the
coordinates. The pair is one rule because the spelling is the source's rather
than the model's: polars and parquet write a hole as a null, pandas has only
NaN, and None in a pandas column is NaN by the time either lane sees it.
Sparsity is the absent row.
Refused¶
| a declared parameter with no data | names the parameter |
| a source nothing can be read as a table from | names the shapes that are read |
an xr.DataArray |
names to_series(), this being a lane's output rather than an input |
a dims: [] parameter whose source has more than one row |
one value broadcast everywhere has one row |
| a dict or a sequence for a parameter over more than one dim | each runs along one dimension |
| a sequence whose length is not the dimension's | positional, so one entry per label |
| a sequence for a dimension nothing else supplies labels for | names the three ways to supply them |
| a key naming neither a parameter nor a dimension | names the near miss |
a table missing a declared dim column, or value |
names the columns needed |
a value column carrying a null or a NaN |
names the parameter and the coordinates |
| a label outside the dimension's index | names the parameter and the strays |
| two rows for one coordinate | |
| a lookup with two values for one label | |
| a lookup value that is not a label of its target | |
| a dimension carrying lookups with no index | |
| a dimension nothing can supply labels for | names both ways to fix it |
| a dimension the file declares and the caller also supplies | names the declaration and the colliding key |
| a lookup whose map the file declares and the caller also supplies | names the map and the colliding column |
| a declared map whose labels nothing supplies | names the map, and asks only for the labels |
| a declared map keyed by something the labels do not carry | names the lookup and the strays |
a column that is not the declared dtype |
names both, and the declaration the data would satisfy |
| a divisor with no value where the model divides by it | names the parameter and how many rows (absence) |
| a comparison's whole constant side with no value where the row is built | the same, naming the constraint |
| a bound parameter with no value where the variable exists | names both models the two repairs build |
Accepted¶
| an undeclared column in a table | ignored |
| a coordinate with no row | sparse data gives sparse variables; what a missing row means where it is read is absence. diagnostics().sparse_parameters says which parameters arrived short of their dims, so a lost row is at least visible (api) |
| a value that is readable and wrong | bound as given; no number is second-guessed |
The index is what makes a stray label a stray¶
A dimension whose labels came from the parameters instead would read that as a third generator, and answer a different question.
Growing or replacing the data¶
A model that is already built takes new numbers with
rebind, and a sweep over slices of
one dimension is solve_over. Both bind through the rules
above.
The opt-in linopy lane binds by these same rules, refusals included.