Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
36 changes: 18 additions & 18 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,30 +13,39 @@ and this project adheres to [Semantic Versioning](http://semver.org/spec/v2.0.0.
without loading them into memory; `--engine` chooses the engine explicitly.
- `vec merge --crs first` keeps the CRS of the first dataset, the default is still EPSG:4326.
- `merge_parquet` accepts `properties` to restrict the merged properties.
- `vec merge --no-strict` warns instead of failing if the merged dataset would be invalid
and writes it anyway; `merge_parquet` has a `strict` parameter (default: `True`).
See the README for the differences between the modes.

### Changed

- **Breaking:** `vec merge` keeps all properties by default, `--include` restricts them to
the core properties plus the given ones and `--exclude` removes any property. It warns if
a required property is not included.
the core properties plus the given ones and `--exclude` removes any property. It reports
required properties that are not included, also collection-only ones.
- **Breaking:** `vec merge` is strict by default: it fails if the merged dataset would be
invalid, e.g. because an id repeats within a collection.
- The message for missing required values names the missing properties with counts,
instead of showing the SQL condition.

### Fixed

- `merge_parquet` no longer drops a property that each part kept as a constant in
its collection but on which the parts disagree; it becomes a column again, as in
`vec merge`.
its collection but on which the parts disagree; it becomes a column again, with the
data type of its schema, as in `vec merge`.
- Merging no longer overwrites the schemas of a collection that occurs in multiple
datasets, it unites them.
datasets, it unites them and rejects two versions of the same schema in one collection.
- Merging warns about collection-only properties that differ between the datasets
and have to be removed.
- Merging fills in missing collection values of datasets that only list a single
collection in `schemas`.
- Merging hydrates array and object constants correctly, instead of spreading them
over the rows, dropping them (`vec merge`) or failing (`merge_parquet`).
- `vec merge` handles constants correctly: arrays and objects are no longer spread over
the rows or dropped, constants that all datasets share stay in the collection, and
constants get the data type of their schema, e.g. binary values are decoded instead
of written as base64 text.
- Merging reports a constant that doesn't fit the data type of its schema, with the
property and the file, instead of failing with an Arrow error.
- `merge_parquet` checks and numbers repeating ids per collection instead of across
all collections, and checks the required properties per collection.
- `merge_parquet` hydrates date-time, date and binary constants with their types and
NaN as null.
- `merge_parquet` recomputes the bbox, which was empty for the rows of GeoParquet 1.0 parts.
- Columns without any value are no longer moved to the collection, which wrote NaN
(invalid JSON) or failed for nullable integers.
Expand All @@ -46,15 +55,6 @@ and this project adheres to [Semantic Versioning](http://semver.org/spec/v2.0.0.
- Validation checks the schemas of all collections, not only the first one.
- Validation no longer fails with a `KeyError` for GeoParquet files with multiple
collections, but without a collection column.
- The in-memory merge keeps constants that all datasets share in the collection, instead
of failing on or moving array and object constants into the rows.
- `vec merge` with DuckDB keeps rows without a required value (with a warning) and
rows with an empty geometry, like the in-memory merge.
- `merge_parquet` warns about a constant that doesn't fit the type of its schema.
- Merging rejects two versions of the same schema in one collection.
- `vec merge --exclude` no longer reads datasets other than GeoParquet twice.
- Merging warns when `--include` drops a required collection-only property.
- Merging accepts features again whose collection can't be determined.

## [v0.3.1] - 2026-09-24

Expand Down
27 changes: 27 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -132,6 +132,33 @@ Local GeoParquet files that are all in the target CRS are merged with DuckDB, so
All other datasets (e.g. GeoJSON or datasets that need to be reprojected) are merged in memory.
Use `--engine` to choose the engine explicitly.

`-i` and `-e` apply to all properties, including collection-level metadata.
Constants that differ between the datasets are moved from the collection metadata to the features.

#### Strict and non-strict mode

By default, `vec merge` is strict: problems that make the merged dataset invalid are errors
and no dataset is written.
With `--no-strict`, `vec merge` is fail-safe: these problems are reported as warnings and
the dataset is written anyway. Check it with `vec validate` afterwards.

| Problem | Strict (default) | `--no-strict` |
| ------- | ---------------- | ------------- |
| A required property has no value for some features | Error | Warning, the property is written as nullable |
| An id repeats within a collection | Error | Warning |
| The collection of a feature can't be determined | Error | Warning, the collection is left empty |
| `-i` or `-e` removes a required property | Error | Warning |
| A required collection-only property differs between the datasets | Error | Warning, the property is removed |
| An optional collection-only property differs between the datasets | Warning, the property is removed | Warning, the property is removed |
| A feature has an empty or missing geometry | Error | Warning, the feature is kept |
| A collection-level value doesn't fit the data type of its schema | Error | Warning, the value is left empty |
| A collection implements multiple versions of a schema, e.g. of an extension | Error | Error |
| The schemas of the datasets conflict | Error | Error |

The strict mode only checks what a merge can break or check with little effort.
It doesn't validate the values against the schemas, e.g. patterns or value ranges,
so a merged dataset is only valid if the values of the source datasets are valid.

Check `vec merge --help` for more details.

### Create JSON Schema from Vecorel Schema
Expand Down
13 changes: 10 additions & 3 deletions tests/test_convert_duckdb.py
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
import json
import re
import sys

import geopandas as gpd
Expand Down Expand Up @@ -730,10 +731,16 @@ def test_constants_that_do_not_fit_their_type_are_reported(value, dtype, capsys)
log = Converter()
logger.remove()
logger.add(sys.stdout, format="{message}", level="DEBUG", colorize=False)
table = _constants_table({"x": value}, {"x": {"type": dtype}}, log=log)
table = _constants_table({"x": value}, {"x": {"type": dtype}}, log=log, source="part.parquet")

assert table.column("x").to_pylist() == [value]
assert f"'x' doesn't fit its type {dtype}" in capsys.readouterr().out
# the value would fail the writer, so it's left empty
assert table.column("x").to_pylist() == [None]
message = f"x: Value {value!r} doesn't fit data type {dtype}"
out = capsys.readouterr().out
assert message in out and "(in part.parquet)" in out

with pytest.raises(ValueError, match=re.escape(message)):
_constants_table({"x": value}, {"x": {"type": dtype}}, log=log, strict=True)


def test_merge_parquet_checks_ids_that_convert_generated(tmp_folder, capsys):
Expand Down
Loading
Loading