Add nest() and unnest() for nested tables - #128
Merged
Merged
Conversation
parallel::detectCores() shells out on Linux and costs about 4 ms, and every verb calls bt_default_threads(). Cache the detected value for the session; the basetable.threads option is still read on every call.
The row-index conversion called INTEGER() once per row, so an indexed subset of 1e6 rows took about 37 ms. Hoisting the pointer brings it to about 5 ms.
Single integer or logical keys fell through to the generic byte-key hash map. Use a value-indexed table when the value range is small and an int hash map otherwise, and hoist the id copy in bt_group_id_. Grouping 1e6 integers drops from about 220 ms to about 12 ms.
nest() gathers each group's columns in one native pass (numeric columns on worker threads). unnest() computes element lengths and stacks the nested frames natively, falling back to the general fill-and-coerce path when the frames differ.
# Conflicts: # NEWS.md
ATTRIB() is no longer part of the R API in R 4.6, which broke the build on current R. Attribute presence now uses ANY_ATTRIB() (R >= 4.5, falling back to ATTRIB() on older R), and unnest() compares column attributes through reusable zero-length carriers filled by Rf_copyMostAttrib().
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
nest()andunnest()for nested tables, native kernels for both, and a benchmark against tidyr and dplyr.nest(data, by, name = "data")collapses each group into a data frame held in a list-column, in first-appearance order. NA keys form their own group.unnest(data, cols)expands list-columns of data frames or vectors back into rows and drops empty elements.unnest()falls back to the general fill-and-coerce path when the nested frames differ.inst/benchmarks/benchmark-nest.Rand a README section compare againsttidyranddplyrat 1e6 rows. basetable is faster in all six scenarios (10, 1e3, 1e5 groups).Two general performance fixes found while profiling are separate commits in this branch and could be split into their own PR:
bt_default_threads()cachesdetectCores()(about 4 ms per verb call).row_index()no longer callsINTEGER()per row, and single integer or logical keys are grouped with a dense table.Checks: full test suite passes,
R CMD check: Status OK.Note:
NAMESPACEis not roxygen-generated in this repo, so the exports anduseDynLibentries were edited by hand.