So in some cases, when we do a query, the results of one part of the query are innately shared-across-nodes; for instance, a hypothetical Distinct query. More generally, we allow cross-index queries; calls can have "index=foo" in them. This patch lets us handle that without duplicating that query all over. Before we actually start doing the separate calls, we run the query once from the coordinating node, then patch the results in, and send relevant subsets over to each client, etcetera. Also provides slightly friendlier (and I hope faster) support for converting bitmaps to/from sets of rows. We also add an extension interface, and some fancy stuff to let us define new calls, which use this. They're sort of tied together because the first extension I wanted to implement needed precomputed calls. The extension API lets us create extensions using `pkg/plugin` (with all its associated limitations, unfortunately), then query them at load time for functionality. This also implies some revamping of the argument validation for PQL, like verifying that functions exist and knowing things about their argument types. So basically this is an overly intrusive patch, and would be better as separate patches, but they're hard to detangle. add trivial execution-time profiling What if you could ?profile=true on a query and get some numbers back? That'd be really cool. We already have tracing/spans, but right now, those only generate any data if you have something set up for them to trace to. Add a fancy wrapper that lets us generate our own tracing data, and dump it into the request response, if ?profile=true. add a sample extension, add missing features to extension interface Implement a naive probabilistic filter extension as an example of what an extension looks like. In the process, discover multiple omissions in the bitmap API. Well, I did *say* it was experimental. |
||
|---|---|---|
| .. | ||
| testdata | ||
| btree.go | ||
| btree_test.go | ||
| container_stash.go | ||
| containers_btree.go | ||
| containers_slice.go | ||
| containers_test.go | ||
| fuzz_test.go | ||
| fuzzer.go | ||
| generation_debug.go | ||
| generation_nodebug.go | ||
| inst.go | ||
| naive.go | ||
| naive_test.go | ||
| nop_inst.go | ||
| README.md | ||
| roaring.go | ||
| roaring_helpers_test.go | ||
| roaring_internal_test.go | ||
| roaring_nop_paranoia.go | ||
| roaring_nop_sentinel.go | ||
| roaring_nop_stats.go | ||
| roaring_paranoia.go | ||
| roaring_sentinel.go | ||
| roaring_stats.go | ||
| roaring_test.go | ||
| source.go | ||
| unmarshal_binary.go | ||
The Fuzzer
For complete documentation on go-fuzz, please see: https://github.com/dvyukov/go-fuzz
The fuzzer in relation to the roaring package checks the Bitmap.UnmarshalBinary function found in roaring.go. In order to use the fuzzer, you can follow these steps:
cd $GOPATH/src/github.com/pilosa/pilosa/roaring
go-fuzz-build ./
You must now make the workdir/corpus directory. This is achieved by:
mkdir workdir/corpus
The fuzzer needs some input to start the fuzzing with. Copy some sample Pilosa fragments into the workdir/corpus folder. For example:
cp ~/.pilosa/my-index/my-field/views/standard/fragments/0 workdir/corpus
Once you have copied your sample inputs, you are ready to run the fuzzer:
go-fuzz -bin=roaring-fuzz.zip -workdir=workdir -func=FuzzBitmapUnmarshalBinary
Understanding the Fuzzer Output
The fuzzer will output something similar to the follwoing:
2015/04/25 12:39:53 workers: 8, corpus: 124 (12s ago), crashers: 37, restarts: 1/15, execs: 35342 (2941/sec), cover: 403, uptime: 12s
The most important part of the output is the crashers and cover. The crashers records how many combinations were discovered that fail and the cover tells you how much code is being accessed.
For a complete explanation of the output, please see: https://github.com/dvyukov/go-fuzz.
The fuzzer will document the crashers in a folder labeled "crashers." It will record the fragment and the error that was produced in two separate files within this folder. This is the final product.
Happy Fuzzing!