Repository navigation
Lasso not working on segmentationLayer #494
Description
Activity
Hi @ArneDefauw, this is currently the expected behavior, unfortunately. The lasso currently works by determining which (x, y) points associated with each cell/spot/observation are within the lasso polygon boundaries. For spots or embedding points, this is easy, because each spot is defined by a single (x, y) coordinate.
For segmentations, either polygons or bitmasks, it is less clear how to map each to an individual (x, y) point. To work around this, we created the obsLocations data type, so that you can provide a single (x, y) point (such as a centroid, or any arbitrary coordinate) corresponding to each segmentation.
Internally, the Spatial view constructs a quadTree using these points, to enable the lasso functionality:
For spatialdata.zarr specifically, we need to define a way to load the data (e.g., from the annotating table element, like
sdata.tables[some_table].obsm['X_spatial'], and/or from specified X/Y columns of the polygons Parquet dataframe) and then we need to supportobsLocationsas part of the configurationoptionsfield - see this TODO statementOnce supporting this on the JS side, we would need to do a little bit more to support this on the Python side at
vitessce-python/src/vitessce/wrappers.py
Line 1410 in 9faf440
def __init__(self, sdata_path: Optional[str] = None, sdata_url: Optional[str] = None, sdata_store: Optional[Union[str, zarr.storage.StoreLike]] = None, sdata_artifact: Optional[ln.Artifact] = None, image_path: Optional[str] = None, region: Optional[str] = None, coordinate_system: Optional[str] = None, obs_spots_path: Optional[str] = None, obs_segmentations_path: Optional[str] = None, obs_points_path: Optional[str] = None, obs_points_feature_index_column: Optional[str] = None, obs_points_morton_code_column: Optional[str] = None, table_path: str = "tables/table", is_zip=None, coordination_values=None, **kwargs):
In the meantime, to provide
obsLocationscoordinates in the absence of the SpatialData support, you can provide the coordinates via a CSV file (example) or an AnnData object (example). In Python, you can useobs_locations_pathparameter of the AnnDataWrapper class or you can use the CsvWrapper class.
You will just need to ensure that theobsTypeprovided viacoordination_valuesmatches with your obsSegmentations data (e.g., bothcell), as Vitessce uses this to associate the information coming from different files to each other https://github.com/vitessce/vitessce/blob/5bcf8d7351e14b29f477f228e8f472bc35ffade0/packages/vit-s/src/data-hooks-multilevel.js#L274
In the long run, to make it easier for users, we should ideally not require specifying the extra explicit (x,y) point correspondences, as this adds unnecessary work. Instead, we should just compute the polygon or bitmask overlaps. However this would be more work to implement. We have an issue to track this at vitessce/vitessce#1956
Hi @keller-mark thanks for the reaction, and the detailed explanation!
lasso on centroids via obsLocations specified via the AnnDataWrapper or CsvWrapper is a good solution for most scenario's I believe.
I tried your suggested solution on the spatialdata blobs dataset, but lasso in the spatial view still generates empty selections.
I used below minimal example:
import tempfile from os.path import join import numpy as np import pandas as pd import spatialdata path = tempfile.mkdtemp(prefix="spatialdata_blobs") sdata = spatialdata.datasets.blobs() sdata # add a dummy categorical column to convince ourself that table is linked to segmentation mask labels = np.random.choice(["A", "B", "C"], size=sdata["table"].obs.shape[0]) sdata["table"].obs["dummy_category"] = pd.Categorical( labels, categories=["A", "B", "C"] ) spatialdata_filepath = join(path, "blobs.spatialdata.zarr") """ # centroids obtained by using harpy import harpy as hp aggregator = hp.utils.RasterAggregator( mask_dask_array=sdata[ "blobs_labels"].data[ None, ... ], image_dask_array=None, ) df=aggregator.center_of_mass() centroids = df[ df[ "cell_ID" ]!=0 ][ [ 2,1 ] ].values """ centroids = np.array( [ [330.09258152, 78.35026898], [264.0552962, 378.97834836], [150.49335749, 235.56662641], [68.01902229, 359.18572305], [242.69662363, 201.91236346], [458.47912317, 218.07124217], [477.66862483, 358.46728972], [135.1428051, 422.40400729], [423.101002, 284.59599198], [85.39643005, 36.7592362], [67.0419056, 230.18129687], [452.63431877, 26.44280206], [19.89012426, 393.87965991], [154.5, 97.0], [473.0218509, 495.86118252], [157.96740548, 493.74445893], [43.03916449, 496.54177546], [162.0078637, 13.20445609], [367.01190476, 497.65873016], [155.22014388, 329.6057554], [21.4254062, 143.79911374], [103.48511905, 473.55208333], [301.84532925, 239.86983155], [357.90895062, 223.89197531], [252.01866252, 99.18662519], [77.93396226, 154.02201258], ] ) sdata["table"].obsm["spatial"] = centroids sdata.write(spatialdata_filepath, overwrite=True)and created the vitessce config as follows:
import os from IPython.display import HTML, display from vitessce import ( AnnDataWrapper, VitessceConfig, get_initial_coordination_scope_prefix, ) from vitessce import ( CoordinationLevel as CL, ) from vitessce import ( CoordinationType as ct, ) from vitessce import ( ViewType as vt, ) from vitessce.wrappers import ObsSegmentationsOmeZarrWrapper vc = VitessceConfig( schema_version="1.0.18", name="ObsLocations + OME-Zarr + ObsSets (dummy)", ) dataset = vc.add_dataset("cells") dataset.add_object( AnnDataWrapper( adata_path=os.path.join(sdata.path, "tables/table"), obs_feature_matrix_path="X", obs_locations_path="obsm/spatial", obs_set_paths=["obs/dummy_category"], obs_set_names=["Cluster"], coordination_values={ "obsType": "cell", "fileUid": "segmentation", }, ) ) dataset.add_object( ObsSegmentationsOmeZarrWrapper( img_path=os.path.join(sdata.path, "labels/blobs_labels"), coordination_values={ "obsType": "cell", "fileUid": "segmentation", }, ) ) spatial = vc.add_view("spatialBeta", dataset=dataset) layer_ctrl = vc.add_view("layerControllerBeta", dataset=dataset) obs_sets = vc.add_view(vt.OBS_SETS, dataset=dataset) ( obs_type, obs_color, obs_set_sel, obs_highlight, obs_set_color, obs_set_highlight, ) = vc.add_coordination( ct.OBS_TYPE, ct.OBS_COLOR_ENCODING, ct.OBS_SET_SELECTION, ct.OBS_HIGHLIGHT, ct.OBS_SET_COLOR, ct.OBS_SET_HIGHLIGHT, ) obs_type.set_value("cell") obs_color.set_value("cellSetSelection") obs_set_sel.set_value(None) obs_highlight.set_value(None) obs_set_color.set_value(None) obs_set_highlight.set_value(None) spatial.use_coordination( obs_type, obs_color, obs_set_sel, obs_highlight, obs_set_color, obs_set_highlight, ) obs_sets.use_coordination( obs_type, obs_set_sel, obs_color, obs_highlight, obs_set_color, obs_set_highlight, ) vc.link_views_by_dict( [spatial, layer_ctrl], { "segmentationLayer": CL( [ { "fileUid": "segmentation", "segmentationChannel": CL( [ { "obsType": obs_type, "obsColorEncoding": obs_color, "obsSetSelection": obs_set_sel, "obsHighlight": obs_highlight, "obsSetColor": obs_set_color, "obsSetHighlight": obs_set_highlight, "spatialTargetC": 0, "spatialChannelOpacity": 0.75, } ] ), } ] ) }, scope_prefix=get_initial_coordination_scope_prefix("A", "obsSegmentations"), ) vc.layout(spatial | (layer_ctrl / obs_sets)) url = vc.web_app() display(HTML(f'<a href="{url}" target="_blank">Open in Vitessce</a>'))Thanks!
Tiny comment on this. We are considering adding centroids by default for each dataset parsed with
spatialdata-io(and store them inadata.obsm['spatial']). Also we are considering defining a "canonical"/"normalized"/"simple" (not sure what's a good name)SpatialDataobject which guarantees that the centroids are present.If we introduce this it would be possible to rely on the centroids for operations.
Hi @LucaMarconato that would be nice!
I think we should be careful when adding these centroids to adata.obsm["spatial"] in spatialdata-io, because it will not be always clear for users (and for downstream applications like vitessce), in what coordinate system these centroids are defined.For example if a table annotates labels (raster) (e.g. for visium HD, see https://github.com/vitessce/vitessce-python/blob/main/docs/notebooks/spatial_data_visium_hd.ipynb). In that case the labels are a grid of around 4000*4000, with an affine transformations defined on it ( in order to align it with the H&E.)
In this case, centroids of the spatial element that is annotated by the table are meaningless, unless we also specify transformations (for example in .uns["spatialdata_attrs"]). Otherwise Vitessce will have no idea how to render these centroids (now they are rendered 'as is', and users of vitessce need to be careful to 'apply' the necessary transformations on the centroids, before providing them to vitessce).
Or maybe you will assume that the centroids 'inherit' the transformations from the spatial element that is annotated by the table?You are correct in bring this up, we will indeed allow to assign a coordinate system to a table (or a transformation to a coordinate system if we implement this before being compatible with NGFF 0.6).
Reacted by Arne Defauw
First of all, thanks for the great package!
I want to do lasso selection on a segmentationLayer, but using the provided example on the spatialdata blobs data, this did not work, live example here and here for CODEX data.
Using a spotLayer, e.g. for visiumHD (live example here ), this does work as expected.
Strange enough, when I add a path to an embedding via obs_embedding_paths to the SpatialDataWrapper (or AnnDataWrapper), and do lasso in the scatterplot (e.g. umap), this does translate to a cell selection spatially.
Is this expected behaviour, and is it currently not possible to do lasso on a segmentationLayer?
Happy to provide more details, thanks!