Commit graph

366 commits

Author SHA1 Message Date
Matt Jaffee
e469285fe3
add fragment mmap tracking and limiting
in the case that the map limit is reached, we'll fall back to reading the file
into memory normally.
2019-03-19 12:55:26 -05:00
Todd Gruben
327aa70924
add failure path for mmap 2019-03-19 12:55:25 -05:00
Seebs
054cb206d5 improve union-related benchmarking
Add a benchmark to test a specific case where UnionInPlace is
underperforming the naive union operation badly.

Also, the UnionBulk test was reusing a bitmap, meaning that it ended
up doing a lot of unions into a bitmap that already had all the
bits it was supposed to have. This broke a couple of other tests
in unexpected ways.

We also now use UnionInPlace in importRoaring, and test it
in the container combinations tests via a wrapper.
2019-03-14 15:16:49 -05:00
Matt Jaffee
a28141c466
revert to Union for importRoaring
UnionInPlace is still heavily affected by
https://github.com/pilosa/pilosa/issues/1875 where containers that exist in an
incoming bitmap can cause massive unnecessary allocations of bitmap containers
when an array of short length is all that's needed.
2019-03-11 17:43:55 -05:00
Matt Jaffee
52d43fb4e2
add smallPath for importRoaring
this converts the rowSet to a map from a slice which might be bad... benchmarks
will tell.
2019-03-11 17:43:55 -05:00
Matt Jaffee
19807ff3a7
use num containers to decide which direction to union
avoids doing a potentially expensive f.storage.Count()
2019-03-11 17:43:55 -05:00
Matt Jaffee
e33ca2d0ae
use UnionInPlace in import-roaring
get the count of the existing fragment and compare it to the incoming bits to
decide which should be unioned into the other. This should generally result in
far fewer allocations, though there is much work that needs to be done within
UnionInPlace to further improve things.

unrelatedly, I added a TODO to change the long-query-time option to move it out
of cluster. It should probably be happening at the API level so that different
handlers can reuse it, but if we're going to do that we'll want to make sure
that any potentially time intensive operations are pulled into api from
handler (e.g. protobuf decoding)
2019-03-11 17:43:55 -05:00
Matt Jaffee
e54dbd6731
simplify row/lastRow comparison in bulkImport 2019-03-05 12:20:21 -06:00
Matt Jaffee
7af64e382c
comments to make import mutex less confusing 2019-03-04 21:38:09 -06:00
Matt Jaffee
8476fffaa7
always operate on storage in importPositions regardless of smallWrite
This greatly simplifies the code, and with the recent addition of DirectAddN and
DirectRemoveN should be as or more performant than doing the separate bitmap and
union (in most cases, unsorted data could still be slower). Perhaps more
importantly, it is also less allocation heavy than the union approach. Also
makes it trivial to get the counts of changed bits, so I've cleaned up the stats
to show number of bits we're importing/clearing and the number of bits that
actually changed.
2019-03-04 21:38:09 -06:00
Matt Jaffee
9b8a97ccb6
maintain column set in bulkImportMutex to guard against repeats 2019-03-04 21:38:08 -06:00
Matt Jaffee
4082ce655a
wip on adding mutex support to random import perf 2019-03-04 21:38:08 -06:00
Matt Jaffee
f06a9f0e6e
positionsForValue appends to existing slices rather than allocating small ones 2019-03-04 21:38:08 -06:00
Matt Jaffee
d429d7c496
code review feedback: add Bitmap.Any and remove unecessary condition 2019-03-04 21:38:07 -06:00
Matt Jaffee
023faebd90
fix bug where opN wasn't getting set/cleared correctly 2019-03-04 21:38:07 -06:00
Matt Jaffee
7037ebf4b3
rename smallPath->smallWrite for consistency 2019-03-04 21:38:07 -06:00
Matt Jaffee
088d618040
increase default MaxOpN 2019-03-04 21:38:07 -06:00
Matt Jaffee
3cbcb238fb
aggregate small bsi imports into a single-write append 2019-03-04 21:38:07 -06:00
Matt Jaffee
ce656bbcda
factor out code to import/clear by positions 2019-03-04 21:38:07 -06:00
Matt Jaffee
537ae99fb9
add import support with aggregated op log writes 2019-03-04 21:38:07 -06:00
Matt Jaffee
0de419c95e
add a SetBit/ClearBit path to bulkImport for small updates
also add benchmarks for this situation and set default MaxOpN higher which
benchmarks suggest is a good idea
2019-03-04 21:38:06 -06:00
Matt Jaffee
daa87d8e12
fix staticcheck warnings 2019-01-21 14:24:11 -06:00
Travis Turner
c66daabc81
convert the anti-entropy logic to use ImportRoaring instead of QueryNode 2018-12-10 21:05:08 -06:00
Ben Johnson
8e49332b25 Add distributed tracing. 2018-11-21 15:08:33 -06:00
Matt Jaffee
8bc1104585
fix fragment checksums race condition 2018-11-20 14:16:23 -06:00
Matt Jaffee
65f478470f
logging cleanup - start with lowercase unless reporting error or warning 2018-11-20 14:08:06 -06:00
Seebs
a203313143 move Logger and Stats to their own packages
I'd like to add stat tracking to Roaring, which means it
has to be able to import the stats package, which means
stats has to be a package rather than part of the pilosa
package. If stats stops being in pilosa, it still needs
a way to import logger, so logger also has to leave the
pilosa package. Then everything using them needs to import
them and use package selectors on their names.

This doesn't actually add the stats support to roaring,
it just makes it so there's a way to import the stats
code from something in the roaring package.
2018-11-15 15:10:44 -06:00
Yuce Tekol
70f85211d9
prevent panic in Bitmap.UnmarshalBinary when there is no data 2018-11-15 22:06:21 +03:00
Matt Jaffee
30590c83bd
Merge branch 'master' into new-rows-iterate 2018-10-25 08:14:58 -05:00
Travis Turner
5fb6ef224c
add clear support for ImportRoaring 2018-10-23 17:37:57 -05:00
Travis Turner
a7a15c64a2
support clear imports to int fields. fix bug in fragment.sum 2018-10-23 17:37:57 -05:00
Travis Turner
caf8e06712
add clear functional option for imports 2018-10-23 17:37:57 -05:00
Matt Jaffee
b6a953cade
cleanup groupby - more comments, remove panic, remove dup test 2018-10-16 19:56:54 -05:00
Matt Jaffee
fb706ab883
add GroupBy Rows(limit) test and fix bug
run all group by tests on two cluster sizes
2018-10-12 18:47:41 -05:00
Matt Jaffee
07d279a155
implement mergeGroupCounts w/o map, remove dead code
move rowFilters to fragment.go

new mergeGroupCounts implementation takes limit into account while merging,
exploits inherent order of group count results.
2018-10-11 19:06:46 -05:00
Matt Jaffee
9d896c5d2f
implement alternate groupByIterator using fragment rowIterator
doesn't re-intersect the same rows for every record
2018-10-11 18:01:44 -05:00
Matt Jaffee
36a539d24e
Merge branch 'master' into new-rows-iterate 2018-10-09 12:27:05 -05:00
Matt Jaffee
1474884f5d
use existing var instead of recalculating
silly mistake - thanks todd
2018-10-09 10:27:57 -05:00
Matt Jaffee
e2bbcb28e5
fix linter issues 2018-10-08 19:10:16 -05:00
Matt Jaffee
c172ca0680
combine fragment.rows and rowsForColumn with generalized filter
use filter funcs with closures for state instead of methods on structs. seems a
bit cleaner.
2018-10-08 16:16:39 -05:00
Travis Turner
619bc1bcd9
implements fragment.setRow(row, rowID) 2018-10-05 08:45:05 -05:00
Matt Jaffee
8cd82af2e7
remove extraneous fragment.rows* methods
variadic filters makes separate methods unnecessary
2018-10-02 09:31:13 -05:00
Matt Jaffee
4f3f2e1a49
remove noFilter and filterWithOffsetLimit
can use an empty list of filters and a list of offsetFilter followed by limit
filter respectively
2018-10-02 09:25:07 -05:00
Matt Jaffee
ed1b09a1cd
fix columnID<>shard checks in executor and fragment
fragment panics if rowsForColumn is called with a column id not in the
fragment's shard. The justification for this is that we're wasting resources if
we're sending requests for a specific column to any shard other than the one
which contains that column.
2018-09-28 10:17:53 -05:00
Matt Jaffee
78ff75690d
convert Rows to use previous/limit
pass previous+1 directly to fragment.rows so that the iterator can seek directly
to the start point. handle limit inside reduce so it can skip out early and
avoid extra allocation.
2018-09-27 16:14:40 -05:00
Travis Turner
4a411067a5
remove Set() (and Clear()) from vector interface, and add error to return 2018-09-24 16:16:22 -05:00
Travis Turner
b6e99734f0
implement ClearRow() query 2018-09-24 15:52:43 -05:00
Travis Turner
4f17dbdbf1
add support for Bool fields
prevent import of non-boolean row values to bool fields
2018-09-21 09:25:41 -05:00
Travis Turner
995a24d0af
ensure mutex imports unset previous columns 2018-09-20 11:25:43 -05:00
Matt Jaffee
54ce537327
copy non roaring-import code from Todds's row-iterate PR
tests passing
2018-09-17 16:01:51 -05:00