Commit graph

5199 commits

Author SHA1 Message Date
Seebs
aab42e97a4 tune b+tree values
Did some benchmarking with b+tree values. The actual interactions
appear to be slightly inconsistent; some values seem to help more
in cases with higher OpN in benchmarks, others with lower OpN. It
appears that the practical consideration may be what happens
when a snapshot gets triggered; smaller kx/kd appear to reduce
costs there, but increase costs between snapshots. This is a
bit of guesswork.

Numbers are slighly under powers of 2, because that means that the
total actual sizes of k and d end up fitting nicely in alloc
pool sizes.
2019-03-28 15:42:00 -05:00
Matthew Jaffee
71e7f62909
Merge pull request #1917 from jaffee/perf-regression
fix importRoaring perf regression [no changelog]
2019-03-28 15:35:08 -05:00
Matt Jaffee
1aafd95adc
make arg naming consistent 2019-03-28 15:11:08 -05:00
Matt Jaffee
c651ff9299
use BTree bitmap in importRoaring
sliceContainers very slow to union into
2019-03-28 13:49:01 -05:00
Matt Jaffee
77a0b6b353
add pathological import benchmark 2019-03-28 13:49:00 -05:00
Matthew Jaffee
5ee49904b5
Merge pull request #1915 from jaffee/benchmarking-tweaks
run fewer concurrency level benchmarks, add bench Makefile target
2019-03-28 12:03:29 -05:00
Matt Jaffee
e9db8eb2c2
run fewer concurrency level benchmarks, add bench Makefile target
The benchmarks take an absurdly long time to run, and I think these are the
largest offenders. Dropping to two concurrency cases 2 and 16 should give a
pretty good idea.
2019-03-27 13:32:56 -05:00
seebs
4955dff22f
Merge pull request #1859 from seebs/seebs/serverinfo
add server stats to /info endpoint
2019-03-26 11:10:37 -05:00
Seebs
5ee87bbf7c add server stats to /info endpoint
report the approximate hardware specs (CPU speed, cores, memory)
of the server in the /info endpoint. This may be useful when
benchmarking.

We do some workarounds because gopsutil's core count output is
confusingly different between Linux and Darwin, and the MHz output
is usually wrong on Linux. Intel's app notes say to just parse
the model string. Whyyyyyyy.
2019-03-26 09:14:51 -05:00
Matthew Jaffee
5e0413eac1
Merge pull request #1911 from jaffee/importValue-data-race
test concurrent value imports, fix race
2019-03-25 16:23:17 -05:00
Matt Jaffee
714f89c65c
simplify locking in importValue
may be a slight perf cost, but the simplicity is well worth it
2019-03-25 14:27:05 -05:00
Matt Jaffee
4420d72196
test concurrent value imports, fix race 2019-03-25 14:25:22 -05:00
Matthew Jaffee
9fda9cf6a3
Merge pull request #1910 from jaffee/profiling-stuff
implement config options for block profile rate and mutex fraction
2019-03-25 14:24:36 -05:00
Cody Soyland
52062c0a27
linkify SetMutexProfileFraction in docs
Co-Authored-By: jaffee <matthew.jaffee@gmail.com>
2019-03-25 13:45:43 -05:00
Cody Soyland
54a6e0ef84
linkify SetBlockProfileRate in docs
Co-Authored-By: jaffee <matthew.jaffee@gmail.com>
2019-03-25 13:45:28 -05:00
Matt Jaffee
9f4a9421bc
fix missing quote in toml tag. unclear how test could pass without it 2019-03-25 12:16:24 -05:00
Matt Jaffee
599b2f4a9e
implement config options for block profile rate and mutex fraction
set sane defaults. The performance overhead seems to be negligible, and this will allow us to obtain mutex and blocking profiles from running Pilosas by default.
2019-03-25 11:38:01 -05:00
Matthew Jaffee
b031b45cbe
Merge pull request #1906 from jaffee/1905-close-files
implement global open file counter using syswrap
2019-03-25 09:53:38 -05:00
Matt Jaffee
53dfa9b7f2
remove rename of columnIDs and add comment 2019-03-23 14:52:16 -05:00
Matt Jaffee
e7f65cf7be
implement global open file counter using syswrap
close files after using them if global max is passed.

I originally implemented this without the global count—just always closing files
when done with them, and reopening for new writes. This was crazy slow for that
one test that uses mustSetBits in a big loop. I modified the test to use
importRoaring and everything worked better (though much more slowly).

After adding the global counter, I ran the tests with that one test using
mustSetBits again, and the performance was similar to master. After completing
this PR, I ran the tests with the max limit set to 5—they still passed but were
much slower.
2019-03-23 14:52:16 -05:00
Cody Soyland
b24bc8bb4b
Merge pull request #1909 from codysoyland/golang-1.12
Add Go 1.12 to CircleCI
2019-03-22 20:30:15 -05:00
Cody Soyland
3096202980 Make workflow require Go 1.12, not Go 1.11 2019-03-22 17:20:08 -05:00
Cody Soyland
38af019113 Default to Go 1.12 2019-03-22 17:06:17 -05:00
Cody Soyland
20137986e5 Add Go 1.12 to CircleCI 2019-03-22 16:59:20 -05:00
seebs
0678c539a1
Merge pull request #1901 from seebs/seebs/smallc
make Containers smaller, especially when they have small contents
2019-03-22 16:56:44 -05:00
Seebs
35593f99df move comment to right place 2019-03-22 16:31:29 -05:00
Seebs
bdbd9c1f47 add missing BCE slices in intersection 2019-03-22 16:31:29 -05:00
Seebs
cd81a9a33f hint to the bounds checker for bitmapRepair
You might wonder why `i <= bitmapN-4`. Answer: The compiler isn't
smart enough for the stride analysis to figure out that `i <= bitmapN`
actually guarantees that. If you set the limit to something not a
multiple of stride, though, it can't figure out *anything* about
things. But for some reason, `i < bitmapN-3` fails badly (it
actually adds bounds checks not present with `i < bitmapN`), but
`i <= bitmapN - 4` works.

This reduces runtime of bitmapRepair by about 14%.
2019-03-22 16:31:29 -05:00
Seebs
8475b97d87 set cap more carefully on unsafe slices
Treating a pointer as a pointer to a large array of bytes,
or converting back the other way, isn't totally insane, but
it does create slices with an extremely large cap. This bit
me while I was trying to build the 16-byte packed Container
structure, but it's probably actually worth fixing in general.
2019-03-22 16:31:29 -05:00
Seebs
8041785ea4 clean up some leftover bits from previous implementation
It used to be useful/desireable to set the other slices to nil when
setting a new slice type, it's no longer useful, take some of those
out.

Also reuse the already-computed run count when converting arrays
and bitmaps to runs.
2019-03-22 16:31:29 -05:00
Seebs
2af5d64e2c unbreak a subtle breakage that only test cases could hit
It turns out the logic for "don't update everything if
the incoming slice pointer is the stash" is wrong; it should
really be "don't update everything if the incoming slice
pointer is the one we already have".

The reason this breaks is that one of the tests directly
sets the mapped bit. This breaks my assumption that we'd
never be using the stash and have the mapped bit set, and
that in turn breaks my assumption that the pointer
of an incoming array can't be the stash address unless
we were previously using the stash. If unmap moved us
to non-stashed memory, then a future write could try to
write, notice that it would fit in the stash, copy the
data ... and not update the pointer because the stash
pointer was handled separately.

This way, if you do that, you can end up not using the
stash when you probably could, but you get the expected
behavior. But also, don't set the mapped bit directly.
(I guess there's a good reason to for the test case,
which is using it to verify that unmaps happen when
modifications happen.)

Also the unmap functions should indicate that they have
successfully unmapped, which may help performance in
some test cases.
2019-03-22 16:31:29 -05:00
Seebs
5c8106bead drop slice implementation
The actually-a-slice implementation of Container was useful in
debugging but does not spark joy.
2019-03-22 16:31:29 -05:00
Seebs
45e8978835 messing around with the performance of unmap
Noticed in profiling that unmap wasn't being inlined. Also noticed
that every call is on a specific container type, so now they're
specialized and small enough to inline.
2019-03-22 16:31:29 -05:00
Seebs
117942c0f3 Add stash-based implementation of Container
This implementation, controlled by the build flag "container24s",
is similar to the single-slice container implementation, but goes
a bit further. First, instead of using a native slice as its internal
storage, it uses pointer/len/cap as distinct values, and only int32
ranges for len and cap. Second, it has a small region of additional
storage which it uses as a backing store by default for arrays or
runs. The idea is that, if you request a new empty array container,
you get one with a pre-allocated virtual slice big enough for five
values, actually stored in the Container. This is useful because
Go's allocator has size classes for 16 and 32 bytes, and the
Container comes in at 24 bytes worth of storage -- meaning that if
you allocate a container, you're allocating 32 bytes anyway, so we
might as well use that space to avoid extra allocations.

This includes some test fixups because DeepEqual was testing
too much equality in some tests.

Also, we simplify unionArrayArray to postpone creating a Container
until we're ready.
2019-03-22 16:31:29 -05:00
Seebs
47dcb5b4a7 Abstract away access to container slices
On a 64-bit machine, the slices in a Container consume 72
bytes, and the Container itself is 80. But we only use one
slice at a time! This patch shifts us to keeping a single
slice in the Container, and converting provided slices to
and from that type when we want to update it. (It is not
safe to access the slice through the wrong type.)

We also add some new tests, conditional on a build tag
called `roaringparanoia`. These tests will be optimized
away entirely by the compiler when the tag isn't
present, because the conditionals use a const. These catch
possible errors like trying to access the bitmap slice
of a non-bitmap container.

We also eliminate all direct creation of Container literals,
so we can mess with the internals more. (On reflection
and study, we decided not to go to the fancier design where
references to .n and .typ were also converted to function
calls, which would have allowed packing those attributes
more tightly, because it was a lot more overhead and a lot
of work to keep track of.)

There's some circumstances where we appear to have been
relying on incorrect guesses about the nature of containers.
For instance, in xorBitmapRun, there's logic that makes sense
only if the output's a run container, but it's not, it's a
bitmap container. This creates strange behavior sometimes,
though. Several of these are corrected now.
2019-03-22 16:31:29 -05:00
Matthew Jaffee
86ea040639
Merge pull request #1908 from jaffee/bitmap-any-quick-fix
quick fix for Bitmap.Any bug [no changelog]
2019-03-21 16:09:59 -05:00
Matt Jaffee
d202e6a1a0
quick fix for Bitmap.Any bug
Want to make empty containers a thing of the past, but that can wait for another
day.
2019-03-21 15:46:19 -05:00
Matthew Jaffee
7550b5445a
Merge pull request #1900 from jaffee/validate-shard
Validate shard
2019-03-20 23:06:22 -05:00
Matt Jaffee
33b54c68d5
add lock on cluster.OwnsShard 2019-03-20 22:04:20 -05:00
Matt Jaffee
97ba8771bc
add test for only opening owned shards 2019-03-20 22:04:20 -05:00
Todd Gruben
418a8788ed
gofmt missing 2019-03-20 22:04:20 -05:00
Todd Gruben
186f034b16
missed commit 2019-03-20 22:04:20 -05:00
Todd Gruben
8edd2b3d13
applied travis suggestions 2019-03-20 22:04:20 -05:00
Todd Gruben
27492a11cc
some formating issues 2019-03-20 22:04:19 -05:00
Todd Gruben
38de65eac0
only load shards that are applicable to node 2019-03-20 22:04:19 -05:00
Matthew Jaffee
66d750bd53
Merge pull request #1904 from jaffee/mmap-fail-docs
add config docs for max-map-count
2019-03-19 20:04:51 -05:00
Matt Jaffee
a7e77566a4
add config docs for max-map-count 2019-03-19 16:21:24 -05:00
Matthew Jaffee
f299473658
Merge pull request #1903 from jaffee/mmap-fail
Mmap fail
2019-03-19 14:02:14 -05:00
Matt Jaffee
226f15446b
lock MaxMapCount and fix unused var 2019-03-19 12:55:26 -05:00
Matt Jaffee
e469285fe3
add fragment mmap tracking and limiting
in the case that the map limit is reached, we'll fall back to reading the file
into memory normally.
2019-03-19 12:55:26 -05:00