Skip to content

summarize step: high peak memory from merging full tables #1124

Description

@i-am-sijia

Summary

The summarize step can use a lot of memory on large models. Three things contribute:

  • Wide merge. It merges trips with tours_merged using every column of both tables, although the summarize specs usually reference only a few of them.
  • Merged inputs. It takes persons, persons_merged, households, households_merged and tours_merged as step inputs. The framework therefore builds the full merged tables before the step body runs, and persons is held twice (once on its own and once inside persons_merged).
  • Temporaries. Every summarize.csv row whose Output starts with _ stays in locals_d until the step ends. Many large temporaries therefore all stay in memory at the same time.

Observed

  • For the Temporaries, a user reordering summarize.csv so each temporary is deleted right after its last use did not change the peak—deleting temporaries cannot lower a peak that is set earlier in the Wide Merge. RSS sampling can also miss a short peak.

Run 1: Memory usage during summarize step, original version without any memory reduction helpers.

Image

Run 2: Delete each Temporary right after its last use.

Image

We can see a small reduction near the end of Run 2, which reflects the impact of deleting Temporaries. However, the peak memory did not reduce. The previous spikes happen during the Wide Merge.

Proposed Changes

  1. _del directive. An Output of _del in summarize.csv removes the named temporaries (comma-separated) from the expression namespace. The code change for this has been implemented in Implement deletion of temporary variables in summarize.py #1062. A unit test to confirm the deletion has been added in Implement deletion of temporary variables in summarize.py (reviewed) #1116.
  2. Trim before merging. Drop columns from trips and tours_merged that no spec references, using util.drop_unused_columns. Column references come from summarize.csv, the preprocessor spec and the BIN/AGGREGATE settings. Unsuffixed names are kept for any _trip/_tour reference so the merged column names don't change. Controlled by a new DROP_UNUSED_COLUMNS setting (default True) and skipped when EXPORT_PIPELINE_TABLES is True. A test checks that the output CSVs are identical with and without trimming.

Drafted code changes in #1125. Tested the memory reduction for Boston, see Run 3 below.

Run 3: Drop unused variables before wide merge of trips and tours

Image
  1. Possible follow-up (not implemented). Take the base tables as step inputs, select only the needed columns, and join them inside the step. This would avoid building the full persons_merged, households_merged and tours_merged tables at all. It needs a way to know which source table each referenced column comes from, and the other users of persons_merged (for example shadow_pricing) must keep working.

  2. Chunking? Chunking with some additional engineering for summarize may help memory, but sometimes summaries do need to be calculated on the full dataset, otherwise the denominators would be incorrect.

Limitations

  • Columns referenced through names built at runtime cannot be detected.
  • The input tables are already loaded when the step starts, so trimming reduces the merge and later steps' memory, not the inputs themselves.
  • Memory benefit depends on the data and the contents in the summarize.csv (how many summaries are being created).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions