Compare commits

...

150 commits

Author SHA1 Message Date
Travis Turner
9015c00da9
Merge pull request #84 from travisturner/foreign-index
Add FieldOption.ForeignIndex
2020-01-10 15:45:02 -06:00
Travis
a6a2f84bd5 During Holder.Open, apply foreign index after all indexes open
In the case where a field with a foreign index opens before the
foreign index has opened (and is available as a reference in the
holder), push the field into a queue to have its foreign index
applied once all indexes have opened.
2020-01-10 12:28:28 -06:00
Travis
3d3286a9ca fix an issue caused by empty column list defaulting to IDs 2020-01-10 12:28:28 -06:00
Travis
f79fde43e3 exclude internal fields (i.e. _exists) from Inspect output 2020-01-10 12:28:28 -06:00
Travis
742135dc10 Get ForeignIndex string value when reading BSI field.
In the `Inspect` function in `server/grpc.go`, getting
the value of an `int` field with a foreign index to
an index with `Keys()`, we need to return the string
key value instead of the BSI int value for the field.

This commit also changes the method `Field.keys()` to be
exported as `Field.Keys()` so that it's accessible in
the server package.
2020-01-10 12:28:28 -06:00
Travis
35f9dfa374 remove write portion of extension data race 2020-01-10 12:28:28 -06:00
Travis
881d3bef06 Adjust the FieldOption logic to be in place prior to field.Open().
This commit changes the order of FieldOption application so that
it's always set before field.Open() is called.

This was required because field.Open() now uses some of the values
from FieldOptions to determine if/when to use a particular
translateStore. For example, when FieldOptions.ForeignIndex is set,
the translateStore from the foreign index is retrieved during
field.Open().
2020-01-10 12:28:28 -06:00
Matt Jaffee
4d8f307c5e handle string values in ImportValueRequest sorting 2020-01-10 12:28:28 -06:00
Matt Jaffee
132cf7cc1c add StringValues to proto ImportValueRequest, update proto versions
I ran:

brew upgrade protobuf
GO111MODULE=off go get -u github.com/gogo/protobuf/protoc-gen-gofast

I'm not sure if everything is still going to work, but I'm excited to
find out!
2020-01-10 12:28:27 -06:00
Travis
b22d0143e4 loadNewExtensions is unused, but included for completeness 2020-01-10 12:28:27 -06:00
Travis
1542cbefc0 Add FieldOption.ForeignIndex
This allows a BSI field to have an option indicating
that it is a foreign key to another index. If the foreign
index has column keys, then this field handles string values
by using the foreign index's translate store.
2020-01-10 12:28:27 -06:00
Matthew Jaffee
7aef0743e1
Merge pull request #86 from jaffee/field-char-limit
allow field and index names up to 230 characters
2020-01-09 14:43:55 -06:00
Matt Jaffee
0a1de81441
allow field and index names up to 230 characters
The lowest limitation I've seen on any filesystem we care about is 255
characters. 230 leaves enough space that an index or field could be
backed up and have a timestamp and file extension appended while
still allowing for much longer index and field names.
2020-01-09 13:02:22 -06:00
tgruben
457789effd
Merge pull request #62 from tgruben/clearvalue
add support to clear value for column
2020-01-07 13:12:21 -06:00
Todd Gruben
643884aeb3 refactored clear to fix q2;removed unused comment 2020-01-07 12:55:32 -06:00
Todd Gruben
26e3460413 removed view creation from ClearValue 2020-01-07 10:55:45 -06:00
Todd Gruben
d843959904 fixed comments; removed create fragment 2020-01-07 10:39:56 -06:00
Todd Gruben
75e017cf4f Merge remote-tracking branch 'upstream/enterprise' into clearvalue 2020-01-07 10:28:47 -06:00
Cody Soyland
29f7448b89
Merge pull request #83 from codysoyland/expvar-compatibility
Initialize expvar lazily to prevent panic if importing both Pilosa v1 and v2
2020-01-02 17:56:42 -06:00
Cody Soyland
8d32005ede Initialize expvar lazily to prevent panic if importing both Pilosa v1 and v2. 2020-01-02 15:43:34 -06:00
Travis Turner
5643afac47
Merge pull request #82 from travisturner/row-response-sorter
WIP: RowResponseSorter for sorting a list of RowResponse based on sort params
2019-12-31 13:18:05 -06:00
Travis
0fffd9a0cb RowResponseSorter for sorting a list of RowResponse based on sort paraters 2019-12-31 11:52:01 -06:00
Travis Turner
c247805d96
Merge pull request #81 from travisturner/inspect-field-output
Remove empty field check in Inspect()
2019-12-27 13:42:36 -06:00
Travis Turner
98df5672e9
Merge branch 'enterprise' into inspect-field-output 2019-12-27 13:26:38 -06:00
Travis Turner
768de9dc3a
Merge pull request #80 from travisturner/row-response-error
Add StatusError to RowResponse for better error handling.
2019-12-27 13:25:38 -06:00
Travis
3be4141382 Return the correct data type label in grpc header
Based on the pilosa field type, return the correct data type
label in the gRPC column header.
2019-12-27 12:43:45 -06:00
Travis
73090e05c8 Add StatusError to RowResponse for better error handling.
This PR adds a `StatusError` to the `pproto.RowResponse` type, which
allows a stream to pass an error on the stream (encoded into
the `RowResponse.StatusError`). This can be checked downstream
for matching `EOF` or `err != nil` and handled appropriately.

This is helpful mainly with the `RowResponse` reducers which run in
goroutines. Instead of trying to manage a separate channel of errors
from those goroutines, we just follow the grpc model and send the
error with the stream.
2019-12-27 12:43:45 -06:00
Travis Turner
94f97297eb
Merge pull request #77 from travisturner/all-shard
Allow All() to be called at the shard level
2019-12-27 12:35:41 -06:00
Travis
d34e38f134 Remove empty field check in Inspect()
The check for field existence is not necessary; since we
add the `_id` field to every response then at the very
least that field will be returned.

This check was preventin a query like `select _id from ...`
from returning any results.
2019-12-26 22:27:46 -06:00
Travis
586a13e942 Allow All() to be called at the shard level 2019-12-20 22:44:46 -06:00
Cody Soyland
7b30b91448
Merge pull request #74 from codysoyland/docker-build-vendor
Vendor modules before building docker image so private modules can be downloaded
2019-12-20 17:52:26 -06:00
Cody Soyland
d4117f3137 Vendor modules before building docker image so private modules can be downloaded 2019-12-20 16:05:18 -06:00
seebs
47316d9f2f
Merge pull request #70 from seebs/unionAA
Simplify unionArrayArray
2019-12-20 13:49:51 -06:00
seebs
d4d3d75e28
Merge branch 'enterprise' into unionAA 2019-12-20 13:28:57 -06:00
Matthew Jaffee
d86a3c3f2f
Merge pull request #55 from seebs/pluginfix
Pluginfix
2019-12-20 12:53:09 -06:00
Matt Jaffee
fe0f57651e
build with distinct by default 2019-12-20 12:21:47 -06:00
Cody Soyland
49b2029656
Run "go mod vendor" outside of Docker so authenticated modules may use system credentials 2019-12-20 12:21:47 -06:00
Seebs
7e1fd8392f
go.mod/go.sum changes for using molecula/ext
This pins us to the initial external release of molecula/ext, which
with any luck will be the only one. (Narrator: It was not to be the
only one.) We also use GOPRIVATE so we don't need a replace directive.
2019-12-20 12:21:47 -06:00
Seebs
0eba050054
stop using pkg/plugin, start using build tags
After a few experiments with pkg/plugin, I'm ready to concede that the
people warning me it was unsuitable for production use were in fact
correct.

In the brave new world, the "ext" package is moved to its own module
outside pilosa. This means that importing it doesn't imply any need to
version-check against pilosa; we can just use versioned copies of the
ext package, which can be public because it doesn't contain anything
we need to care about keeping proprietary.

Then we can, conditional on build tags, import modules from a
neighboring repo which contains the actual implementations, and if
they're imported, their init functions register them.
2019-12-20 12:21:47 -06:00
Seebs
f51c2dbc42
use extensions through build tags 2019-12-20 12:21:47 -06:00
Travis Turner
179fb910f7
Merge pull request #71 from travisturner/all-limit-offset
Add All() support to PQL, including limit and offset
2019-12-18 22:10:42 -06:00
Travis
361e51cb41 Add All() support to PQL, including limit and offset
This PR is meant to get all columns from an index
based on the TrackExistence row.

`All()` is a PQL function that can be used as a typical
row object. Optional arguments are `limit` and `offset`.
2019-12-18 18:00:15 -06:00
Seebs
83aa505673 Simplify unionArrayArray
Also short-circuit it in some cases.
2019-12-17 15:30:56 -06:00
Travis Turner
1af016df84
Merge pull request #68 from travisturner/linter-fix
Fix an impossible code path raised by the linter
2019-12-17 09:04:37 -06:00
Travis
532caa0fbf fix an impossible code path raised by the linter 2019-12-16 21:56:57 -06:00
Travis Turner
6f556fb880
Merge pull request #63 from travisturner/row-field-label
Wrap return types: RowIdentifiers, Pair, and []Pair
2019-12-16 07:42:20 -06:00
Travis
4f7f4f58b1 add field to SignedRow, and implement its grpc response 2019-12-14 15:56:06 -06:00
Travis
3b7b54094a update clustertests to use v2 (and go 1.13) 2019-12-13 18:45:43 -06:00
Travis
5cb37834a0 Wrap return types: RowIdentifiers, Pair, and []Pair
This PR adds a field name (string) to the return types
which represent the values from a specific field. For example,
a TopN query on field `x` would be `TopN(x)` and have results
like:
```
[]Pair{
  {ID: 14, Count: 10},
  {ID: 3, Count: 8},
  {ID: 7, Count: 3},
}
```
In order to know what field this result type refers to, we wrap
`[]Pair` in a new struct called `PairsField` which contains an
addition `Field` string where `x` is stored.

This is useful for informing the gRPC server how to construct
more appropriate headers for the result stream (in this case,
the column headers can now be "x" and "count").

Similar logic was applied to `RowIdentifiers` and `Pair` as well.
2019-12-13 15:36:51 -06:00
Travis Turner
c5aeed0715
Merge pull request #57 from tgruben/fix-minmax-count
return total match counts for either min or max
2019-12-13 15:28:27 -06:00
Todd Gruben
4523a4d693
removed uneeded test run 2019-12-13 15:10:28 -06:00
Todd Gruben
16171b3e65
return total match counts for either min or max 2019-12-13 15:10:28 -06:00
Todd Gruben
95f2864abe applied suggestions by @travisturner 2019-12-11 12:13:57 -06:00
Todd Gruben
7a2d90ade8 add support to clear value for column 2019-12-09 16:28:03 -06:00
Travis Turner
c4e339f72b
Merge pull request #59 from travisturner/not-found-code
Add grpc NotFound code where applicable
2019-12-03 17:44:30 -06:00
Travis
a5652a182c add grpc NotFound code where applicable 2019-12-03 11:53:20 -06:00
Matthew Jaffee
521ea603d0
Merge pull request #52 from molecula/import-col-attrs
Import col attrs
2019-12-02 01:25:08 -06:00
Matt Jaffee
87ee83f4cb
check errors in test to fix lint 2019-12-01 07:33:25 -06:00
Matt Jaffee
9e3029b969
re run generate-protoc 2019-12-01 07:33:24 -06:00
Alan Bernstein
e470b3276e
Add generated proto 2019-12-01 07:33:24 -06:00
Alan Bernstein
cf7d668b49
Support import column attrs in client 2019-12-01 07:33:24 -06:00
Alan Bernstein
c3e8284f6c
Test for presence of column attrs 2019-12-01 07:33:24 -06:00
Alan Bernstein
628def3db7
Clarify some error messages 2019-12-01 07:33:24 -06:00
Alan Bernstein
aba67364b1
Add support for importing column attrs 2019-12-01 07:33:23 -06:00
Travis Turner
76bb3985f0
Merge pull request #53 from travisturner/having-between
Add support for BETWEEN type conditions in the having clause.
2019-11-30 21:58:45 -06:00
Travis
52debbc389 Add support for BETWEEN type conditions in the having clause.
There is a TODO in the `StringWithSubj` method because the value
types really depend on the subject type (for example, `count` uses
uint64, while `sum` uses int64). I'm waiting to address this
until we decide how to handle sums of floats (Decimal), because
that will affect this logic as well.
2019-11-30 11:59:09 -06:00
Travis Turner
0247a9073c
Merge pull request #51 from travisturner/groupby-having
Add "having" support to GroupBy() queries
2019-11-29 23:01:15 -06:00
Travis
24d02c1920 Add "having" support to GroupBy() queries
This PR adds support for a `having` argument in a `GroupBy` query.
Usage looks like this:
```
GroupBy(Rows(a), having=Condition(count > 10))
GroupBy(Rows(a), aggregate=Sum(field=b), having=Condition(sum > 100))
```
2019-11-28 18:40:57 -06:00
Travis Turner
bc0018b67b
Merge pull request #50 from travisturner/datatype-fixes
Add StreamClient and StreamServer interfaces
2019-11-27 17:21:55 -06:00
Travis
3bc0dc28f0 Add StreamClient and StreamServer interfaces
In order to standardize results as streams of RowResponse,
this PR introduces two interfaces `StreamClient` and
`StreamServer`) which mirror the grpc stream interfaces.
Upstream users (sqlmapper, vdsm, etc) can implement
instances of these interfaces to ensure that results can
stream through the entire sytem in an expected way.

This PR also fixes a couple of missing data types.
2019-11-27 16:10:46 -06:00
Matthew Jaffee
210676c927
Merge pull request #39 from molecula/groupby-sum
Groupby sum
2019-11-27 16:04:42 -06:00
Matt Jaffee
114c1a9df8
remove unecessary span from executor tracing 2019-11-27 11:09:06 -06:00
Matt Jaffee
60397d3e8f
add more tracing around groupBy 2019-11-27 11:00:13 -06:00
Matt Jaffee
f2f9ea01dc
get group by aggregates working with PQL validation and grpc streaming 2019-11-27 11:00:13 -06:00
Ben Johnson
ea9914dba0
Add optional GroupBy() 'aggregate' field.
This commit adds an `aggregate` field that allows a `Sum()` call
to be executed for every returned group.
2019-11-27 11:00:13 -06:00
Matt Jaffee
81a4d32cdd
allow IDs to be passed even when keys enabled
This change allows one to query Pilosa fields and indexes directly
with integer row and column ids even when key translation is
enabled. This was previously disallowed during query
translation... I'm not sure why, but it can be quite useful for
debugging and testing to be able to use IDs directly. I have a test in
go-pilosa which uses this functionality.

I also simplified a bunch of the test code which was of the form:
```
else {
    if blah {
    }
}
```

to be:

```
else if blah {
```

which I think is pretty harmless.

I also changed a snapshot log line that has been bugging me to be
Debug level so that it isn't generating lots of useless logs for long
running Pilosa instances.
2019-11-27 08:41:40 -06:00
tgruben
16d5db97a6
Merge pull request #47 from tgruben/inspect-fix
fixed bug in slice container seek
2019-11-26 11:35:51 -06:00
Todd Gruben
ca0247e35b fixed bug in slice container seek 2019-11-26 09:21:52 -06:00
seebs
407c309640
Merge pull request #46 from seebs/distinctfix
Distinctfix
2019-11-22 17:29:24 -06:00
Seebs
d79ecbad86 recompute shards for cross-index calls
It turns out that we need to recompute the set of shards whenever
a query is cross-index. Otherwise we get partial results in unexpected
ways sometimes.
2019-11-22 16:02:47 -06:00
Seebs
36ef82d7ac Zero bitmap storage when reusing it for container-as-bitmap
If you don't do this, it works fine the first time you use a given
storage, but after that you start seeing spurious bits.
2019-11-22 16:02:42 -06:00
Seebs
397d93e84b provide an empty filter when a filter was empty 2019-11-22 16:02:31 -06:00
Seebs
26326ac74c recompute shards for cross-index queries
When computing results on another index, recompute list of
shards for that index.
2019-11-22 16:02:26 -06:00
Seebs
6ad39a376e handle precalls and cross-index queries better
There's two actual changes here, but they're closely related.

First, handle named parameters for precalls, not just indexed parameters.
Second, when doing translation for a call, check whether it specifies an
index, and if it does, use that index instead of the current index for
the translation.
2019-11-22 16:02:20 -06:00
seebs
16c3cfa727
Merge pull request #42 from seebs/sqdeadlock
use the right lock for Enqueue
2019-11-18 21:53:00 -06:00
seebs
01309e4fc5
Merge branch 'enterprise' into sqdeadlock 2019-11-18 20:51:40 -06:00
Cody Soyland
d3c8728821
Merge pull request #19 from codysoyland/ci-updates
Update CI to support Pilosa Enterprise (private dependencies)
2019-11-18 13:52:20 -06:00
Cody Soyland
58182ff563 Pilosa Enterprise private CI 2019-11-18 13:30:32 -06:00
Travis Turner
295adbfe67
Merge pull request #44 from travisturner/cache-threshold-deletes
Fixed ranked cache logic to support reducing cached values below the threshold
2019-11-18 12:10:46 -06:00
Travis
ed82a535e5 Fix ranked cache logic to support reducing cached values below
the threshold.

Prior to this commit, if a cache value was reduced to a value
that fell below the threshold, the operation would be ignored
and the cached value would remain at the old, higher value.

This commit also fixes logic which reduces a cached value within
the framework of uint64 values by subracting the absolute value
of the negative value (since adding a negitive doesn't work with
unsigned integers).
2019-11-18 11:39:50 -06:00
Travis Turner
32f754899e
Merge pull request #43 from travisturner/translate-race-in-test
allow for translate store race in test
2019-11-15 16:33:30 -06:00
Travis
0a69fca657 allow for translate store race in test (by using retry)
In this case, the test is reading from the translateStore
replica before the translateStore replication has had time to
deliver its log to the replica. The only way to truly address
this in the translate store would be to route all key misses
that happen on a read-only replica to the primary translate
store (or somehow know when the primary is done sending to
replicas) for actual verification that the key does not exist.
That's more involved than we want to do here; this PR just
addresses the problem in the test.
2019-11-15 07:57:47 -06:00
Seebs
f204b37760 use atomics instead of locking for stats 2019-11-14 17:40:14 -06:00
Seebs
bab077199c use the right lock for Enqueue
The request for a non-read lock blocks until all existing read
locks exit, meaning that if an Immediate operation is already
going for a fragment, an Enqueue operation will hang forever
holding the fragment's lock, while the Immediate operation has
probably relinquished the fragment's lock to wait for the
queue worker to process it. But the queue worker can't process
it, because the incoming Enqueue still holds the fragment's
lock. Solution: Don't block the Enqueue operation like that.
It shouldn't coexist with things that actually change the sq
channels, like Stop(), but it is fine for it to coexist with
other queue operations.
2019-11-14 17:00:34 -06:00
seebs
9cec40e69d
Merge pull request #41 from seebs/roaringProto
Stop using Roaring in protobuf messages until it's supported elsewhere
2019-11-14 12:45:35 -06:00
Seebs
6fc6cc4350 Stop using Roaring in protobuf messages until it's supported elsewhere
go-pilosa uses protobuf to talk to us but doesn't support the roaring
format. Conveniently, there's a kill switch.
2019-11-14 10:54:40 -06:00
Matthew Jaffee
fa9c911860
Merge pull request #37 from travisturner/rename-bool-label
Rename bool label from changed to result
2019-11-13 14:30:14 -06:00
Travis
e8cd48155a
Rename bool lable from changed to result 2019-11-13 14:12:52 -06:00
Matthew Jaffee
4dfeb89b43
Merge pull request #36 from travisturner/includescolumn-keys
Add column keys support to IncludeColumn
2019-11-13 14:09:24 -06:00
Travis
625125bac6 add column keys support to IncludeColumn 2019-11-13 11:34:21 -06:00
Matthew Jaffee
013ee21621
Merge pull request #28 from molecula/pql-float-values
Pql float values
2019-11-13 11:04:36 -06:00
Matt Jaffee
0e445db7ff
improve error message checking field type in import roaring 2019-11-13 10:07:24 -06:00
Travis
26fc621f09
Support integer predicates in Decimal field range queries. 2019-11-13 10:07:24 -06:00
Matt Jaffee
0ac778516e
check error, make linter happy 2019-11-13 10:07:24 -06:00
Matt Jaffee
14dfc9e31d
fix comment typo for BTWN_LTE_LT 2019-11-13 10:07:24 -06:00
Matt Jaffee
9c8ad727b5
allow floats in PQL queries for decimal fields
had to workaround some cruft in the parser that was trying to only
support a BETWEEN query as LTE, LTE. Now we have operations for all
combinations of LT and LTE.

unrelated - changed the port a test was binding to as it conflicted
with a port I was using locally.
2019-11-13 10:07:24 -06:00
Matt Jaffee
0fec16a141
fix integer bug on less than queries.
this was introduced recently to fix another bug. the comment above it
is correct, just the logic was off-by-one. The test shows the issue
and was confirmed to reproduce it and then fix it.
2019-11-13 10:07:23 -06:00
Travis Turner
7401fd1333
Merge pull request #33 from travisturner/grpc-logger
Fixed gRPC server logger; pass logger through from main
2019-11-13 09:56:04 -06:00
Travis Turner
afd1c004f9
Merge branch 'enterprise' into grpc-logger 2019-11-13 08:31:26 -06:00
Travis Turner
d3432478f9
Merge pull request #34 from travisturner/bool-returns
Fix makeRows in gRPC hander to handle a bool result
2019-11-13 08:24:43 -06:00
Travis
7fd5248d98 makeRows in gRPC hander now handles a bool result
This PR adds bool support to the makeRows function
in the gRPC handler.
2019-11-12 16:13:52 -06:00
seebs
1fea1ea375
Merge pull request #22 from seebs/fsckSnapshotExtension
This is a collection of changes that have been pending forever. It improves the snapshot queue performance, adds some amount of recovery for corrupt filles, reduces memory usage in the rowcache, and adds an extension interface. Yes, they should probably have happened separately over time, things happened.
2019-11-12 14:50:34 -06:00
Seebs
1e0873c70b lock BufferLogger for reads/writes
With the new addition of the holder background scan, it's possible
for an open holder to write log messages at arbitrary times. The
TestHolder_Open/ErrIndexName test checks the contents of the output
buffer, but those contents could be changing if the background task
happens to run at the right time. Use trivial locking around that
so that this shouldn't happen.
2019-11-12 12:15:13 -06:00
Seebs
03f3f424aa don't lint PEG files
I was pretty sure I'd done this, but I guess not: Skip linting
the PEG files.
2019-11-12 12:15:13 -06:00
Seebs
8cc7a176b5 license header fixups
Fix up license headers for the extension code, and add the proto
file to the list of things we don't check license headers for.
2019-11-12 12:15:13 -06:00
Seebs
c5136b14db ensmarten snapshot queue
The snapshot queue needs a bit more subtlety. In some cases,
we really do want to do a snapshot right now -- these shouldn't
have to wait for possibly a hundred or more other snapshots
to complete.

In other cases, we don't really care that much whether we do
a snapshot, and just dropping it is probably fine.

To accommodate this, we distinguish between "urgent" and
"normal" snapshots, and between "Immediate" (does an urgent
snapshot, waits for it) and "Enqueue" (might enqueue a snapshot
but *also might not* if we're already busy). There's a
corresponding "Await" to wait for a snapshot, if one is
pending, but not if one isn't.

We also have a background scan that checks the holder. It will
scan pretty actively when it's finding fragments that need
snapshots (no enqueued snapshot, opN > MaxOpN). It pauses
for a second after every hundred fragments that didn't need
snapshots, and for a minute after each holder scan that didn't
find any. So, if you don't need snapshots, it does basically
nothing, if you do, it'll be moderately aggressive about
submitting tasks -- but it always waits if there's *any*
requested snapshots in the queues.

Updates since initial draft:

Check results from Await more consistently, and in one case, use Immediate
instead and then check its error.

Fix a race condition.  The race condition comes about if:

1. You have a limited enough worker pool that this can happen.
(In testing we tend to have a worker pool of 1.)
2. A fragment is in the normal, non-urgent, queue already.
3. An immediate request comes in for that fragment. This always
happens *with the fragment lock held*.
4. A worker thread grabs that fragment from the queue.
5. The worker thread now waits on the lock. Meanwhile, the
immediate request blocks on sending the fragment to the urgent
queue.
6. The worker can't read the urgent queue, and the immediate
request can't send it, so the immediate request can't proceed.

What's supposed to happen is that the immediate request sends
the thing, and gets into Await(), which sleeps on a condition
variable using the lock, which is to say, releases the lock.

The obvious resolution is to let go of the lock, send the
message, and then reclaim the lock. But then we have the
possibility that the message sent ends up with a timestamp
right after a snapshot that happened *after* the Immediate
request was started. Oops. So we create the request, then let
go of the lock, then send the request, then reclaim the lock
and go into the Await state. All is well.

This is on top of more general use of wait groups, etcetera,
to allow us to ensure that any holder scans terminate *before*
we close the channels they might otherwise be trying to write to.
So, shutdown process is now:

* grab lock on queue (workers and scanners don't use the lock)
* mark snapshotqueue done
* wait for holder scans to complete/exit
* close and nil out all the channels
* release lock

Anything trying to submit to this needs to hold the lock, unless
it's a holder scan, so either it got the lock before we did and already
submitted the thing, or it will get the lock after this and not find
a channel to write to; it's just the holder scanner that has an
ongoing thing that might have started a write to the channel *without*
a lock held, because it's expected that it might have to wait minutes
or hours before the write will complete because it's a background task.

Also, rework the background holder scan to grab lists of
indexes/fields/views/fragments, then scan the grabbed/copied lists,
rather than iterating over maps, allowing us to grab the lock when
we're about to access a thing and let it go when done.

There might be a simpler/cleaner way to do this but opinions on how
safe it is are very mixed, so in the mean time, I'm making the range
behavior not depend at all on there being no writes to the various tiers
of holder/index/view/fragment during the background scans.
2019-11-12 12:15:13 -06:00
Seebs
3b696da34a plugins and precomputed data
So in some cases, when we do a query, the results of one
part of the query are innately shared-across-nodes; for
instance, a hypothetical Distinct query. More generally,
we allow cross-index queries; calls can have "index=foo"
in them.

This patch lets us handle that without duplicating that
query all over. Before we actually start doing the
separate calls, we run the query once from the coordinating
node, then patch the results in, and send relevant subsets
over to each client, etcetera. Also provides slightly
friendlier (and I hope faster) support for converting
bitmaps to/from sets of rows.

We also add an extension interface, and some fancy stuff
to let us define new calls, which use this. They're sort
of tied together because the first extension I wanted to
implement needed precomputed calls. The extension API
lets us create extensions using `pkg/plugin` (with all its
associated limitations, unfortunately), then query them
at load time for functionality.

This also implies some revamping of the argument
validation for PQL, like verifying that functions exist
and knowing things about their argument types.

So basically this is an overly intrusive patch, and would
be better as separate patches, but they're hard to detangle.

add trivial execution-time profiling

What if you could ?profile=true on a query and get some
numbers back? That'd be really cool.

We already have tracing/spans, but right now, those only generate
any data if you have something set up for them to trace to. Add a
fancy wrapper that lets us generate our own tracing data, and dump
it into the request response, if ?profile=true.

add a sample extension, add missing features to extension interface

Implement a naive probabilistic filter extension as an example of
what an extension looks like. In the process, discover multiple
omissions in the bitmap API. Well, I did *say* it was experimental.
2019-11-12 12:14:29 -06:00
Seebs
b25eb8f596 Sources and Generations: tracking mmapped files
This code represents an attempt at providing reliable tracking
of whether any bitmaps still in use have access to a given block
of mmapped data, allowing us to unmap the data when nothing is using
it anymore.

The basic approach is as follows: Each mmap is associated with
a new object, called a "generation". A generation reflects
a particular instance of a given file being mapped. When a
bitmap is built from an mmapped data source, the bitmap is
given a pointer to the generation as its Source. When bitmap
operations combine containers from other bitmaps, they
produce new bitmaps that are tagged with the combined set of
sources.

When we snapshot a file, or for some other reason wish to remap
it, the corresponding bitmap has all its containers updated to
use the new storage, and the bitmap's source is changed. However,
previously-handed-out containers might still have references to the
old storage. Those containers would be in bitmaps with the old
source.

After a bunch of study of trying to reference-count and track
this, I realized: We don't actually need to do that, because we
already have something suitable for determining whether anything
can reach a given object. It's the garbage collector.

So we set a finalizer on the generation object, which handles
unmapping. There's additional sanity-checks here to confirm things
like "we thought this generation should be expiring", and we
track timestamps. We could also have things check whether a
given bitmap's source was marked as obsolete "a while ago", but
that isn't implemented yet.

There's a debug version of this which tracks finalization, creation,
and ending timestamps, and has a call to provide diagnostics for
this. Identical generation IDs get separated out with random
suffixes in this case -- there's sometimes a second or third
instance of the same name due to a holder closing and reopening,
but this basically only happens in testing.

Note that generations are still used even when there's no mmapping,
but unless debugging is turned on, they shouldn't propagate much --
we don't consider a generation to be the source of a bitmap unless
the bitmap actually mapped things from that generation's mmapped
storage, or debugging is on.

There's a couple of other, possibly more subtle, changes and
bug fixes that got caught by the testing on this:
* If a fragment is partially opened and then opening some later
  part fails, we close the earlier parts before returning the
  error so we aren't leaving it partially open.
* Several operations on segments which were requesting that a
  frozen copy of a bitmap be created are now actually *replacing*
  their bitmap with the frozen bitmap, rather than discarding it.
* intersectRunRun, if it decides to create an array or bitmap,
  will yield that container instead of discarding it.

And why all of this? Why, so we can actually implement the thing
where when a fragment has a valid roaring bitmap, but the ops log
is corrupt, we can truncate the corrupt part of the ops log and
reopen it. Which I did.

When the generationdebug build tag is in use, every generation
has a finalizer all the time. When it's not, they only get finalizers
when we expect them to be done -- say, when closing a fragment.
This is because finalizers appear to be possibly-expensive.

There's some logical cleanup to openStorage here, dividing part
of its work into applyStorage and importStorage, which have a common
case for handling "there's no data in this file".
2019-11-12 12:14:29 -06:00
Seebs
6654466033 partially implement truncation of fragments for corrupt ops log
Which is to say don't actually implement it, because openStorage
is too messy right now, but this is the rest of the framework,
and now I'm going to digress into fixing openStorage.
2019-11-12 12:14:29 -06:00
Seebs
538768ea9d handle truncated/damaged .available.shards
The available shards file is just a hint to save us a bit
of time later; we don't need it to run and it can get updated
pretty easily later. If we have problems reading it, we
should just report the error, nuke the file, and continue
without it.
2019-11-12 12:14:29 -06:00
Travis
b8f665db1a pass logger through to the grpc server and handler 2019-11-12 11:08:25 -06:00
seebs
1262cd18e0
Merge pull request #30 from seebs/intfixes
Fix an error that could cause imported values to keep high-order bits from previously imported values, and another that could cause BSI fields to store extra bits they don't need.
2019-11-12 00:02:04 -06:00
seebs
b3adf2d4f6
Merge branch 'enterprise' into q2-11-7 2019-11-11 21:39:45 -06:00
Cody Soyland
5194ede82c
Merge pull request #31 from codysoyland/proto-pkg-name
Change proto package name to "pilosa" to not conflict with molecula
2019-11-11 20:51:52 -06:00
Cody Soyland
69e5e523c1
Merge branch 'enterprise' into proto-pkg-name 2019-11-11 17:26:44 -06:00
Seebs
a804a0dfb1 always treat BSI fields as having at least their depth
If you imported only small values, BSI fields could end up
not bothering to clear higher bits in existing values, which
produced strange behaviors.

We also move the computation of requiredDepth, and the change
to the field, down, combining it with the other checks of the
values for min/max being in range.

Without this, a data set with a ludicrously large value in it
could break a BSI field's depth even though the import would then
reject it.
2019-11-11 16:55:26 -06:00
Travis Turner
1a44f02e3c reset fragment.rowCache after importValue 2019-11-11 16:54:51 -06:00
Travis Turner
10514f7ced
Merge pull request #25 from travisturner/includes-column
Add an IncludesColumn() function to PQL
2019-11-11 12:55:23 -06:00
Cody Soyland
b9335c9f5c Change proto package name to "pilosa" to not conflict with molecula. Upgrade protoc to 3.10.1 2019-11-11 12:24:38 -06:00
Travis
ed37ef5dcf Add an IncludesColumn() function to PQL
Usage:
`IncludesColumn(Intersect(Row(a=1), Row(b=2)), column=10)`

The above query will return a `bool` indicating whether the
intersection of rows a-1 and b-2 contains column 10. Because
a single column is specified, this executes on a single shard
(shard=0 in this example).
2019-11-11 08:17:28 -06:00
Travis Turner
198626e657
Merge pull request #29 from travisturner/linter-fixes
fix golangci-lint complaints
2019-11-11 08:06:35 -06:00
Travis
87b8edc4c5 fix golangci-lint complaints 2019-11-10 17:58:46 -06:00
Travis Turner
2a5d79ad83
Merge pull request #26 from travisturner/proto-licence-exception
add proto/pilosa.pb.go to license.exceptions list
2019-11-10 16:31:00 -06:00
Travis
3129b1c841 add proto/pilosa.pb.go to license.exceptions list 2019-11-10 11:33:29 -06:00
Travis Turner
6632821617
Merge pull request #24 from travisturner/cache-size-none
fix cacheSize when cacheType is none (and cacheSize is 0)
2019-11-08 22:30:17 -06:00
Travis
a842dd521c fix cacheSize when cacheType is none (and cacheSize is 0)
There was an edge case where setting cacheType to none
wouldn't zero out its cacheSize. This fixes that edge case.
2019-11-08 15:58:49 -06:00
Matthew Jaffee
d6f2196bf1
Merge pull request #16 from pilosa/tls-grpc
use TLS settings when setting up GRPC server or client
2019-10-30 19:35:56 -05:00
Matt Jaffee
6f21887259
use TLS settings when setting up GRPC server or client 2019-10-30 17:38:48 -05:00
Matthew Jaffee
468cf98811
Merge pull request #15 from pilosa/decimal-to-grpc
add decimal field support to Inspect
2019-10-30 13:33:55 -05:00
Matt Jaffee
a414cada4f
add decimal field support to Inspect
I tested this manually with BloomRPC and curl, but need to write real
tests. Also need to get floats for decimal fields coming out of QueryPQL.
2019-10-30 11:13:48 -05:00
Matthew Jaffee
8f64a4f585
Merge pull request #11 from pilosa/decimal-support
support for decimal fields
2019-10-29 16:49:20 -05:00
Matt Jaffee
a9a4d244ef
fix a bug in the "less than" logic 2019-10-29 16:36:15 -05:00
Matt Jaffee
5dcabfcc7f
support for decimal fields
This commit adds a Decimal field type which is implemented mostly with
the Int field. It adds an optional "Scale" value to the Int field
which means that the values stored in that field are actually meant to
be divided by 10^Scale before being interpreted.

In order to make use of this functionality, we extend the importValue
request to allow a slice of floats rather than just int64. If the
slice of floats is present, each float in the slice is multiplied by
10^Scale and converted to an int64 before being imported. If a slice
of int64 is imported to a Decimal field, it is treated normally, and
scale is ignored. This allows the conversion to be handled at the
client side if desired.

Currently there are Field level methods for querying Float values out
of a decimal field, but no support in PQL or the executor for getting
float values. Going to wait until I can use the generic result type
before doing that, so for now, any values queried will be the scaled
integer values.

needed to add client support for importing float values, and did this
by adding a more general and simplified client method for value
imports.

rewrote api.ImportValue to use the new method which should be more
performant and efficient.

allow floats to be "pilosa import"ed into decimal fields
2019-10-29 16:36:14 -05:00
seebs
4ef7f7e26b
Merge pull request #12 from seebs/profile
Profiling and a couple of minor fixes
2019-10-29 15:24:51 -05:00
Seebs
616ed39771 Skip longest tests when running -short
The cluster timeout/down tests are way more than half the total
time for "go test", and are very unlikely to be of interest in regular
usage, although they matter for CI. Skip them when doing short
tests.
2019-10-29 15:24:09 -05:00
Seebs
7a381f7eaf allow years other than 2017 in licenses
Also clean up the license hash checking a bit. We trim vendor early
in find so we don't have to walk the whole vendor tree only to grep
the files out, and we don't check the license hashes of the exceptions,
and the exceptions are now a plain text file of non-regex strings
we match exactly. Also the license hash code is only written once.

This will help us a lot if development on Pilosa continues through
2018 or later.
2019-10-29 15:24:09 -05:00
Seebs
820c5ce220 add trivial execution-time profiling
What if you could ?profile=true on a query and get some
numbers back? That'd be really cool.

We already have tracing/spans, but right now, those only generate
any data if you have something set up for them to trace to. Add a
fancy wrapper that lets us generate our own tracing data, and dump
it into the request response, if ?profile=true.

We track wall-clock execution time, plus possible arbitrary K/V
pairs. Memory stats are not included, because obtaining them is
surprisingly expensive.
2019-10-29 15:23:37 -05:00
Travis Turner
9a2f5b3b4c
Merge pull request #10 from travisturner/grpc
initial gRPC server implementation
2019-10-29 14:39:16 -05:00
Travis
a7bb90fcd0 initial gRPC server implementation
add makeRows() tests
register the gRPC server
use api.Index() instead of api.Schema()

support most field types in Inspect() query

currently, there's no support for `time` fields.
those will be dependent upon the output format
and the ability to materialize the timestamp from
the time views.

this commit also changes the response type of the
`Inspect()` query to be a tabular `RowResponse`.
2019-10-29 14:27:07 -05:00
Travis
77a81eb2e1 fix bug preventing a Rows() query on a bool field 2019-10-29 14:27:07 -05:00
84 changed files with 17589 additions and 4050 deletions

View file

@ -8,10 +8,13 @@ defaults: &defaults
fast-checkout: &fast-checkout fast-checkout: &fast-checkout
attach_workspace: attach_workspace:
at: . at: .
add-github-auth: &add-github-auth
run: git config --global url."https://moleculacorp:${GITHUB_PERSONAL_ACCESS_TOKEN}@github.com".insteadOf "https://github.com"
jobs: jobs:
setup: setup:
<<: *defaults <<: *defaults
steps: steps:
- *add-github-auth
- checkout - checkout
- restore_cache: - restore_cache:
keys: keys:
@ -28,11 +31,13 @@ jobs:
<<: *defaults <<: *defaults
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: make check-license-headers - run: make check-license-headers
linter: linter:
<<: *defaults <<: *defaults
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: curl -sfL https://install.goreleaser.com/github.com/golangci/golangci-lint.sh | sh -s v1.20.0 - run: curl -sfL https://install.goreleaser.com/github.com/golangci/golangci-lint.sh | sh -s v1.20.0
- run: sudo cp bin/golangci-lint /usr/local/bin/ - run: sudo cp bin/golangci-lint /usr/local/bin/
- run: make golangci-lint - run: make golangci-lint
@ -40,6 +45,7 @@ jobs:
<<: *defaults <<: *defaults
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: make build GOOS=linux GOARCH=arm GOARM=5 - run: make build GOOS=linux GOARCH=arm GOARM=5
- run: make build GOOS=linux GOARCH=arm GOARM=6 - run: make build GOOS=linux GOARCH=arm GOARM=6
- run: make build GOOS=linux GOARCH=arm GOARM=7 - run: make build GOOS=linux GOARCH=arm GOARM=7
@ -48,18 +54,21 @@ jobs:
<<: *defaults <<: *defaults
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: sudo apt-get install lsof - run: sudo apt-get install lsof
- run: make test - run: make test
test-golang-1.13-shard22: test-golang-1.13-shard22:
<<: *defaults <<: *defaults
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: sudo apt-get install lsof - run: sudo apt-get install lsof
- run: make test SHARD_WIDTH=22 - run: make test SHARD_WIDTH=22
test-golang-1.13-race: test-golang-1.13-race:
<<: *defaults <<: *defaults
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: sudo apt-get install lsof - run: sudo apt-get install lsof
- run: - run:
command: make test TESTFLAGS="-race -v -timeout=30m" command: make test TESTFLAGS="-race -v -timeout=30m"
@ -73,6 +82,7 @@ jobs:
<<: *defaults <<: *defaults
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: sudo apt-get install lsof - run: sudo apt-get install lsof
- run: make test ENTERPRISE=1 - run: make test ENTERPRISE=1
test-golang-1.12: test-golang-1.12:
@ -81,6 +91,7 @@ jobs:
- image: circleci/golang:1.12 - image: circleci/golang:1.12
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: sudo apt-get install lsof - run: sudo apt-get install lsof
- run: make test - run: make test
test-golang-1.11: test-golang-1.11:
@ -89,18 +100,21 @@ jobs:
- image: circleci/golang:1.11 - image: circleci/golang:1.11
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: sudo apt-get install lsof - run: sudo apt-get install lsof
- run: make test - run: make test
cluster-tests: cluster-tests:
<<: *defaults <<: *defaults
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- setup_remote_docker - setup_remote_docker
- run: make clustertests-build - run: make clustertests-build
prerelease: prerelease:
<<: *base-test <<: *base-test
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: make prerelease - run: make prerelease
- store_artifacts: - store_artifacts:
path: build path: build
@ -111,6 +125,7 @@ jobs:
<<: *defaults <<: *defaults
steps: steps:
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: make release - run: make release
- store_artifacts: - store_artifacts:
path: build path: build
@ -123,6 +138,7 @@ jobs:
steps: steps:
- run: '[[ -v CIRCLE_PR_NUMBER ]] && circleci step halt || true' # Skip job if this is a PR - run: '[[ -v CIRCLE_PR_NUMBER ]] && circleci step halt || true' # Skip job if this is a PR
- *fast-checkout - *fast-checkout
- *add-github-auth
- run: sudo pip install awscli - run: sudo pip install awscli
- run: make prerelease-upload - run: make prerelease-upload
dockerhub-upload: dockerhub-upload:
@ -130,11 +146,23 @@ jobs:
steps: steps:
- run: '[[ -v CIRCLE_PR_NUMBER ]] && circleci step halt || true' # Skip job if this is a PR - run: '[[ -v CIRCLE_PR_NUMBER ]] && circleci step halt || true' # Skip job if this is a PR
- *fast-checkout - *fast-checkout
- *add-github-auth
- setup_remote_docker - setup_remote_docker
- run: make docker - run: make docker
- run: docker tag pilosa:$(git describe --tags) pilosa/pilosa:master - run: docker tag pilosa:$(git describe --tags) pilosa/pilosa:master
- run: docker login -u $DOCKER_USER -p $DOCKER_PASS - run: docker login -u $DOCKER_USER -p $DOCKER_PASS
- run: docker push pilosa/pilosa:master - run: docker push pilosa/pilosa:master
docker-enterprise-upload:
<<: *defaults
steps:
#- run: '[[ -v CIRCLE_PR_NUMBER ]] && circleci step halt || true' # Skip job if this is a PR
- *fast-checkout
- *add-github-auth
- setup_remote_docker
- run: make docker-enterprise
- run: docker tag pilosa:$(git describe --tags) molecula.azurecr.io/pilosa-enterprise:unstable
- run: docker login -u $DOCKER_USER -p $DOCKER_PASS
- run: docker push molecula.azurecr.io/pilosa-enterprise:unstable
workflows: workflows:
version: 2 version: 2
test: test:
@ -185,11 +213,4 @@ workflows:
only: /^v.*/ only: /^v.*/
branches: branches:
ignore: /.*/ ignore: /.*/
- prerelease-upload:
requires:
- prerelease
- dockerhub-upload:
requires:
- linter
- check-license-headers
- test-golang-1.13

3
.golangci.yml Normal file
View file

@ -0,0 +1,3 @@
run:
skip-files:
- pql/pql.peg.go

View file

@ -1,8 +1,11 @@
FROM golang:1.13.0 as builder FROM golang:1.13.0 as builder
ARG BUILD_FLAGS
ARG MAKE_FLAGS
COPY . pilosa COPY . pilosa
RUN cd pilosa && CGO_ENABLED=0 make install FLAGS="-a" RUN cd pilosa && CGO_ENABLED=0 make install FLAGS="-a -mod=vendor ${BUILD_FLAGS}" ${MAKE_FLAGS}
FROM alpine:3.9.4 FROM alpine:3.9.4

View file

@ -1,17 +1,14 @@
# This Dockerfile is used for cluster testing - it produces a much larger image # This Dockerfile is used for cluster testing - it produces a much larger image
# and includes all of Go as well as some utilities. # and includes all of Go as well as some utilities.
FROM golang:1.11 FROM golang:1.13
LABEL maintainer "dev@pilosa.com" LABEL maintainer "dev@pilosa.com"
COPY . /go/src/github.com/pilosa/pilosa/ COPY . /go/src/github.com/pilosa/pilosa/
RUN cd /go/src/github.com/pilosa/pilosa \ RUN cd /go/src/github.com/pilosa/pilosa \
&& GO111MODULE=on make vendor && CGO_ENABLED=0 make install FLAGS="-a -mod=vendor"
RUN cd /go/src/github.com/pilosa/pilosa \
&& CGO_ENABLED=0 make install FLAGS="-a"
# download pumba for fault injection # download pumba for fault injection
ADD https://github.com/alexei-led/pumba/releases/download/0.6.0/pumba_linux_amd64 /pumba ADD https://github.com/alexei-led/pumba/releases/download/0.6.0/pumba_linux_amd64 /pumba

View file

@ -16,8 +16,16 @@ RELEASE_ENABLED = $(subst 0,,$(RELEASE))
BUILD_TAGS += $(if $(ENTERPRISE_ENABLED),enterprise) BUILD_TAGS += $(if $(ENTERPRISE_ENABLED),enterprise)
BUILD_TAGS += $(if $(RELEASE_ENABLED),release) BUILD_TAGS += $(if $(RELEASE_ENABLED),release)
BUILD_TAGS += shardwidth$(SHARD_WIDTH) BUILD_TAGS += shardwidth$(SHARD_WIDTH)
LICENSE_HASH=$(shell head -13 pilosa.go | shasum | cut -f 1 -d " ") BUILD_TAGS += $(foreach p,$(PLUGINS),plugin$(p))
define LICENSE_HASH_CODE
head -13 $1 | sed -e 's/Copyright 20[0-9][0-9]/Copyright 20XX/g' | shasum | cut -f 1 -d " "
endef
LICENSE_HASH=$(shell $(call LICENSE_HASH_CODE, pilosa.go))
PLUGINS=distinct
export GO111MODULE=on export GO111MODULE=on
export GOPRIVATE=github.com/molecula
export PLUGINS
# Run tests and compile Pilosa # Run tests and compile Pilosa
default: test build default: test build
@ -81,14 +89,14 @@ DOCKER_COMPOSE=internal/clustertests/docker-compose.yml
# running. This will catch changes to internal/clustertests/*.go, but if you # running. This will catch changes to internal/clustertests/*.go, but if you
# make changes to Pilosa, you'll want to run clustertests-build to rebuild the # make changes to Pilosa, you'll want to run clustertests-build to rebuild the
# pilosa image. # pilosa image.
clustertests: clustertests: vendor
docker-compose -f $(DOCKER_COMPOSE) down docker-compose -f $(DOCKER_COMPOSE) down
docker-compose -f $(DOCKER_COMPOSE) build client1 docker-compose -f $(DOCKER_COMPOSE) build client1
docker-compose -f $(DOCKER_COMPOSE) up --exit-code-from=client1 docker-compose -f $(DOCKER_COMPOSE) up --exit-code-from=client1
# Like clustertests, but rebuilds all images. # Like clustertests, but rebuilds all images.
clustertests-build: clustertests-build: vendor
docker-compose -f $(DOCKER_COMPOSE) down docker-compose -f $(DOCKER_COMPOSE) down
docker-compose -f $(DOCKER_COMPOSE) up --exit-code-from=client1 --build docker-compose -f $(DOCKER_COMPOSE) up --exit-code-from=client1 --build
@ -115,14 +123,23 @@ generate-stringer:
generate-pql: require-peg generate-pql: require-peg
cd pql && peg -inline pql.peg && cd .. cd pql && peg -inline pql.peg && cd ..
# dunno if protoc-gen-gofast is actually needed here
generate-proto-grpc: require-protoc require-protoc-gen-gofast
protoc -I proto proto/pilosa.proto --go_out=plugins=grpc:proto
# `go generate` all needed packages # `go generate` all needed packages
generate: generate-protoc generate-stringer generate-pql generate: generate-protoc generate-stringer generate-pql
# Create Docker image from Dockerfile # Create Docker image from Dockerfile
docker: docker: vendor
docker build -t "pilosa:$(VERSION)" . docker build --build-arg BUILD_FLAGS="${FLAGS}" -t "pilosa:$(VERSION)" .
@echo Created docker image: pilosa:$(VERSION) @echo Created docker image: pilosa:$(VERSION)
# Create Docker image from Dockerfile (enterprise)
docker-enterprise: vendor
docker build --build-arg MAKE_FLAGS="ENTERPRISE=1" -t "pilosa-enterprise:$(VERSION)" .
@echo Created docker image: pilosa-enterprise:$(VERSION)
# Compile Pilosa inside Docker container # Compile Pilosa inside Docker container
docker-build: docker-build:
docker run --rm -v $(PWD):/go/src/$(CLONE_URL) -w /go/src/$(CLONE_URL) -e GOOS=$(GOOS) -e GOARCH=$(GOARCH) golang:$(GO_VERSION) go build -tags='$(BUILD_TAGS)' -ldflags $(LDFLAGS) $(FLAGS) $(CLONE_URL)/cmd/pilosa docker run --rm -v $(PWD):/go/src/$(CLONE_URL) -w /go/src/$(CLONE_URL) -e GOOS=$(GOOS) -e GOARCH=$(GOARCH) golang:$(GO_VERSION) go build -tags='$(BUILD_TAGS)' -ldflags $(LDFLAGS) $(FLAGS) $(CLONE_URL)/cmd/pilosa
@ -133,7 +150,7 @@ docker-test:
# Run golangci-lint # Run golangci-lint
golangci-lint: require-golangci-lint golangci-lint: require-golangci-lint
golangci-lint run golangci-lint run --skip-files '.*\.peg\.go'
# Run gometalinter with custom flags # Run gometalinter with custom flags
gometalinter: require-gometalinter vendor gometalinter: require-gometalinter vendor
@ -161,9 +178,8 @@ gometalinter: require-gometalinter vendor
# Verify that all Go files have license header # Verify that all Go files have license header
check-license-headers: SHELL:=/bin/bash check-license-headers: SHELL:=/bin/bash
check-license-headers: check-license-headers:
@! find . -name '*.go' | grep -v '^./vendor' | while read fn;\ @! find . -path ./vendor -prune -o -name '*.go' -print | grep -v -F -f license.exceptions | while read fn;\
do [[ `head -13 $$fn | shasum | cut -f 1 -d " "` == $(LICENSE_HASH) ]] || echo $$fn; done | \ do [[ `$(call LICENSE_HASH_CODE, $$fn)` == $(LICENSE_HASH) ]] || echo $$fn; done | grep '.'
grep -v apimethod_string.go | grep -v pb.go | grep -v peg.go | grep -v lru.go | grep -v btree | grep -v enterprise
###################### ######################
# Build dependencies # # Build dependencies #

166
api.go
View file

@ -23,7 +23,9 @@ import (
"fmt" "fmt"
"io" "io"
"io/ioutil" "io/ioutil"
"math"
"net/url" "net/url"
"sort"
"strconv" "strconv"
"strings" "strings"
"sync" "sync"
@ -146,9 +148,11 @@ func (api *API) Query(ctx context.Context, req *QueryRequest) (QueryResponse, er
} }
execOpts := &execOptions{ execOpts := &execOptions{
Remote: req.Remote, Remote: req.Remote,
Profile: req.Profile,
ExcludeRowAttrs: req.ExcludeRowAttrs, // NOTE: Kept for Pilosa 1.x compat. ExcludeRowAttrs: req.ExcludeRowAttrs, // NOTE: Kept for Pilosa 1.x compat.
ExcludeColumns: req.ExcludeColumns, // NOTE: Kept for Pilosa 1.x compat. ExcludeColumns: req.ExcludeColumns, // NOTE: Kept for Pilosa 1.x compat.
ColumnAttrs: req.ColumnAttrs, // NOTE: Kept for Pilosa 1.x compat. ColumnAttrs: req.ColumnAttrs, // NOTE: Kept for Pilosa 1.x compat.
EmbeddedData: req.EmbeddedData, // precomputed values that needed to be passed with the request
} }
resp, err := api.server.executor.Execute(ctx, req.Index, q, req.Shards, execOpts) resp, err := api.server.executor.Execute(ctx, req.Index, q, req.Shards, execOpts)
if err != nil { if err != nil {
@ -383,7 +387,7 @@ func (api *API) ImportRoaring(ctx context.Context, indexName, fieldName string,
// only set and time fields are supported // only set and time fields are supported
if field.Type() != FieldTypeSet && field.Type() != FieldTypeTime { if field.Type() != FieldTypeSet && field.Type() != FieldTypeTime {
return NewBadRequestError(errors.New("roaring import is only supported for set and time fields")) return NewBadRequestError(errors.Errorf("roaring import is only supported for set and time fields, not '%s' fields.", field.Type()))
} }
errCh := make(chan error, len(nodes)) errCh := make(chan error, len(nodes))
@ -540,7 +544,7 @@ func (api *API) ExportCSV(ctx context.Context, indexName string, fieldName strin
var colStr string var colStr string
var err error var err error
if field.keys() { if field.Keys() {
if rowStr, err = field.translateStore.TranslateID(rowID); err != nil { if rowStr, err = field.translateStore.TranslateID(rowID); err != nil {
return errors.Wrap(err, "translating row") return errors.Wrap(err, "translating row")
} }
@ -889,10 +893,17 @@ func (api *API) FieldAttrDiff(ctx context.Context, indexName string, fieldName s
return attrs, nil return attrs, nil
} }
// ImportOptions holds the options for the API.Import method. // ImportOptions holds the options for the API.Import
// method.
//
// TODO(2.0) we have entirely missed the point of functional options
// by exporting this structure. If it needs to be exported for some
// reason, we should consider not using functional options here which
// just adds complexity.
type ImportOptions struct { type ImportOptions struct {
Clear bool Clear bool
IgnoreKeyCheck bool IgnoreKeyCheck bool
Presorted bool
} }
// ImportOption is a functional option type for API.Import. // ImportOption is a functional option type for API.Import.
@ -916,6 +927,13 @@ func OptImportOptionsIgnoreKeyCheck(b bool) ImportOption {
} }
} }
func OptImportOptionsPresorted(b bool) ImportOption {
return func(o *ImportOptions) error {
o.Presorted = b
return nil
}
}
// Import bulk imports data into a particular index,field,shard. // Import bulk imports data into a particular index,field,shard.
func (api *API) Import(ctx context.Context, req *ImportRequest, opts ...ImportOption) error { func (api *API) Import(ctx context.Context, req *ImportRequest, opts ...ImportOption) error {
span, _ := tracing.StartSpanFromContext(ctx, "API.Import") span, _ := tracing.StartSpanFromContext(ctx, "API.Import")
@ -941,7 +959,7 @@ func (api *API) Import(ctx context.Context, req *ImportRequest, opts ...ImportOp
// check to see if keys need translation. // check to see if keys need translation.
if !options.IgnoreKeyCheck { if !options.IgnoreKeyCheck {
// Translate row keys. // Translate row keys.
if field.keys() { if field.Keys() {
if len(req.RowIDs) != 0 { if len(req.RowIDs) != 0 {
return errors.New("row ids cannot be used because field uses string keys") return errors.New("row ids cannot be used because field uses string keys")
} }
@ -962,7 +980,7 @@ func (api *API) Import(ctx context.Context, req *ImportRequest, opts ...ImportOp
// For translated data, map the columnIDs to shards. If // For translated data, map the columnIDs to shards. If
// this node does not own the shard, forward to the node that does. // this node does not own the shard, forward to the node that does.
if index.Keys() || field.keys() { if index.Keys() || field.Keys() {
m := make(map[uint64][]Bit) m := make(map[uint64][]Bit)
for i, colID := range req.ColumnIDs { for i, colID := range req.ColumnIDs {
@ -1036,6 +1054,10 @@ func (api *API) ImportValue(ctx context.Context, req *ImportValueRequest, opts .
return errors.Wrap(err, "validating api method") return errors.Wrap(err, "validating api method")
} }
if err := req.Validate(); err != nil {
return errors.Wrap(err, "validating import value request")
}
// Set up import options. // Set up import options.
options, err := setUpImportOptions(opts...) options, err := setUpImportOptions(opts...)
if err != nil { if err != nil {
@ -1059,43 +1081,41 @@ func (api *API) ImportValue(ctx context.Context, req *ImportValueRequest, opts .
if req.ColumnIDs, err = index.translateStore.TranslateKeys(req.ColumnKeys); err != nil { if req.ColumnIDs, err = index.translateStore.TranslateKeys(req.ColumnKeys); err != nil {
return errors.Wrap(err, "translating columns") return errors.Wrap(err, "translating columns")
} }
req.Shard = math.MaxUint64
// For translated data, map the columnIDs to shards. If
// this node does not own the shard, forward to the node that does.
m := make(map[uint64][]FieldValue)
for i, colID := range req.ColumnIDs {
shard := colID / ShardWidth
if _, ok := m[shard]; !ok {
m[shard] = make([]FieldValue, 0)
}
m[shard] = append(m[shard], FieldValue{
Value: req.Values[i],
ColumnID: colID,
})
} }
// Signal to the receiving nodes to ignore checking for key translation. // Translate values when the field uses keys (for example, when
opts = append(opts, OptImportOptionsIgnoreKeyCheck(true)) // the field has a ForeignIndex with keys).
if field.Keys() {
var eg errgroup.Group uints, err := field.translateStore.TranslateKeys(req.StringValues)
for shard, vals := range m { if err != nil {
// TODO: if local node owns this shard we don't need to go through the client return errors.Wrap(err, "translating string values")
shard := shard
vals := vals
eg.Go(func() error {
return api.server.defaultClient.ImportValue(ctx, req.Index, req.Field, shard, vals, opts...)
})
} }
return eg.Wait() // Because the BSI field supports negative value, we have to
// convert the slice of uint64 keys to a slice of int64.
ints := make([]int64, len(uints))
for i := range uints {
ints[i] = int64(uints[i])
}
req.Values = ints
} }
} }
// Validate shard ownership. if !options.Presorted {
sort.Sort(req)
}
// if we're importing into a specific shard
if req.Shard != math.MaxUint64 {
// Check that column IDs match the stated shard.
if s1, s2 := req.ColumnIDs[0]/ShardWidth, req.ColumnIDs[len(req.ColumnIDs)-1]/ShardWidth; s1 != s2 && s2 != req.Shard {
return errors.Errorf("shard %d specified, but import spans shards %d to %d", req.Shard, s1, s2)
}
// Validate shard ownership. TODO - we should forward to the
// correct node rather than barfing here.
if err := api.validateShardOwnership(req.Index, req.Shard); err != nil { if err := api.validateShardOwnership(req.Index, req.Shard); err != nil {
return errors.Wrap(err, "validating shard ownership") return errors.Wrap(err, "validating shard ownership")
} }
// Import columnIDs into existence field. // Import columnIDs into existence field.
if !options.Clear { if !options.Clear {
if err := importExistenceColumns(index, req.ColumnIDs); err != nil { if err := importExistenceColumns(index, req.ColumnIDs); err != nil {
@ -1105,11 +1125,89 @@ func (api *API) ImportValue(ctx context.Context, req *ImportValueRequest, opts .
} }
// Import into fragment. // Import into fragment.
if len(req.Values) > 0 {
err = field.importValue(req.ColumnIDs, req.Values, options) err = field.importValue(req.ColumnIDs, req.Values, options)
if err != nil { if err != nil {
api.server.logger.Printf("import error: index=%s, field=%s, shard=%d, columns=%d, err=%s", req.Index, req.Field, req.Shard, len(req.ColumnIDs), err) api.server.logger.Printf("import error: index=%s, field=%s, shard=%d, columns=%d, err=%s", req.Index, req.Field, req.Shard, len(req.ColumnIDs), err)
} }
return errors.Wrap(err, "importing") } else if len(req.FloatValues) > 0 {
err = field.importFloatValue(req.ColumnIDs, req.FloatValues, options)
if err != nil {
api.server.logger.Printf("import error: index=%s, field=%s, shard=%d, columns=%d, err=%s", req.Index, req.Field, req.Shard, len(req.ColumnIDs), err)
}
}
return errors.Wrap(err, "importing value")
}
options.IgnoreKeyCheck = true
start := 0
shard := req.ColumnIDs[0] / ShardWidth
var eg errgroup.Group // TODO make this a pooled errgroup
for i, colID := range req.ColumnIDs {
if colID/ShardWidth != shard {
subreq := &ImportValueRequest{
Index: req.Index,
Field: req.Field,
Shard: shard,
ColumnIDs: req.ColumnIDs[start:i],
}
if req.Values != nil {
subreq.Values = req.Values[start:i]
} else if req.FloatValues != nil {
subreq.FloatValues = req.FloatValues[start:i]
}
eg.Go(func() error {
return api.server.defaultClient.ImportValue2(ctx, subreq, options)
})
start = i
shard = colID / ShardWidth
}
}
subreq := &ImportValueRequest{
Index: req.Index,
Field: req.Field,
Shard: shard,
ColumnIDs: req.ColumnIDs[start:],
}
if req.Values != nil {
subreq.Values = req.Values[start:]
} else if req.FloatValues != nil {
subreq.FloatValues = req.FloatValues[start:]
}
eg.Go(func() error {
// TODO we should elevate the logic for figuring out which
// node(s) to send to into API instead of having those details
// in the client implementation.
return api.server.defaultClient.ImportValue2(ctx, subreq, options)
})
return eg.Wait()
}
func (api *API) ImportColumnAttrs(ctx context.Context, req *ImportColumnAttrsRequest, opts ...ImportOption) error {
span, _ := tracing.StartSpanFromContext(ctx, "API.ImportColumnAttrs")
defer span.Finish()
index, err := api.Index(ctx, req.Index)
if err != nil {
return errors.Wrap(err, "getting index")
}
if err := api.validateShardOwnership(req.Index, uint64(req.Shard)); err != nil {
return errors.Wrap(err, "validating shard ownership")
}
bulkAttrs := make(map[uint64]map[string]interface{})
for n := 0; n < len(req.ColumnIDs); n++ {
bulkAttrs[uint64(req.ColumnIDs[n])] = map[string]interface{}{req.AttrKey: req.AttrVals[n]}
}
if err := index.ColumnAttrStore().SetBulkAttrs(bulkAttrs); err != nil {
api.server.logger.Printf("import error: index=%s, shard=%d, len(columns)=%d, err=%s", req.Index, req.Shard, len(req.ColumnIDs), err)
return errors.Wrap(err, "importing column attrs")
}
return nil
} }
func importExistenceColumns(index *Index, columnIDs []uint64) error { func importExistenceColumns(index *Index, columnIDs []uint64) error {

117
api/client/grpc.go Normal file
View file

@ -0,0 +1,117 @@
// Copyright 2017 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package client
import (
"context"
"crypto/tls"
pb "github.com/pilosa/pilosa/v2/proto"
"github.com/pkg/errors"
"google.golang.org/grpc"
"google.golang.org/grpc/credentials"
)
// GRPCClient is a client for working with the gRPC server.
type GRPCClient struct {
conn *grpc.ClientConn
}
// NewGRPCClient returns a new instance of GRPCClient.
func NewGRPCClient(dialTarget string, tlsConfig *tls.Config) (*GRPCClient, error) {
var opts []grpc.DialOption
if tlsConfig != nil {
creds := credentials.NewTLS(tlsConfig)
opts = append(opts, grpc.WithTransportCredentials(creds))
} else {
opts = append(opts, grpc.WithInsecure())
}
gconn, err := grpc.Dial(dialTarget, opts...)
if err != nil {
return nil, errors.Wrap(err, "creating new grpc client")
}
return &GRPCClient{
conn: gconn,
}, nil
}
// Close closes any connections the client has opened.
func (c *GRPCClient) Close() error {
if c.conn != nil {
return c.conn.Close()
}
return nil
}
// Query returns a stream of RowResponse for the given index and PQL string.
func (c *GRPCClient) Query(ctx context.Context, index string, pql string) (pb.StreamClient, error) {
if c.conn == nil {
return nil, errors.New("client has not established a grpc connection")
}
grpcClient := pb.NewPilosaClient(c.conn)
stream, err := grpcClient.QueryPQL(ctx, &pb.QueryPQLRequest{
Index: index,
Pql: pql,
})
if err != nil {
return nil, errors.Wrap(err, "getting stream")
} else if stream == nil {
return nil, errors.New("could not create stream")
}
return stream, err
}
// Inspect returns a stream of RowResponse for the given index, columns, and filters.
// It is intended to mimic something like "select [fields] from table where recordID IN (...)".
func (c *GRPCClient) Inspect(ctx context.Context, index string, columnIDs []uint64, columnKeys []string, fieldFilters []string, limit, offset uint64) (pb.StreamClient, error) {
if c.conn == nil {
return nil, errors.New("client has not established a grpc connection")
}
if len(columnIDs) > 0 && len(columnKeys) > 0 {
return nil, errors.New("only provide column ids or keys, not both")
}
// Convert columns to proto type IdsOrKeys.
idsOrKeys := &pb.IdsOrKeys{}
if len(columnKeys) > 0 {
idsOrKeys.Type = &pb.IdsOrKeys_Keys{Keys: &pb.StringArray{Vals: columnKeys}}
} else {
idsOrKeys.Type = &pb.IdsOrKeys_Ids{Ids: &pb.Uint64Array{Vals: columnIDs}}
}
grpcClient := pb.NewPilosaClient(c.conn)
stream, err := grpcClient.Inspect(ctx, &pb.InspectRequest{
Index: index,
Columns: idsOrKeys,
FilterFields: fieldFilters,
Limit: limit,
Offset: offset,
})
if err != nil {
return nil, errors.Wrap(err, "getting stream")
} else if stream == nil {
return nil, errors.New("could not create stream")
}
return stream, err
}

View file

@ -19,6 +19,7 @@ import (
"fmt" "fmt"
"math" "math"
"reflect" "reflect"
"strconv"
"strings" "strings"
"testing" "testing"
"time" "time"
@ -30,6 +31,137 @@ import (
"github.com/pilosa/pilosa/v2/test" "github.com/pilosa/pilosa/v2/test"
) )
// attrFun defines a mapping from columnID -> attr value
func attrFun(id uint64) string {
//return fmt.Sprintf("%x", md5.Sum([]byte(strconv.FormatInt(int64(id), 10))))
return strconv.FormatInt(int64(id), 10)
}
func TestAPI_ImportColumnAttrs(t *testing.T) {
/*
columns seconds
100 1.150
1000 1.568
10000 5.156
100000 38.179
*/
c := test.MustRunCluster(t, 2,
[]server.CommandOption{
server.OptCommandServerOptions(
pilosa.OptServerNodeID("node0"),
pilosa.OptServerClusterHasher(&offsetModHasher{}),
)},
[]server.CommandOption{
server.OptCommandServerOptions(
pilosa.OptServerNodeID("node1"),
pilosa.OptServerClusterHasher(&offsetModHasher{}),
)},
)
defer c.Close()
m0 := c[0]
m1 := c[1]
t.Run("ImportColumnAttrs", func(t *testing.T) {
ctx := context.Background()
index := "i"
field := "f"
attrKey := "k"
_, err := m0.API.CreateIndex(ctx, index, pilosa.IndexOptions{})
if err != nil {
t.Fatalf("creating index: %v", err)
}
_, err = m0.API.CreateField(ctx, index, field)
if err != nil {
t.Fatalf("creating field: %v", err)
}
// Generate some attrs for two shards
numAttrs := 100
columnIDs0 := make([]uint64, 0, numAttrs)
attrVals0 := make([]string, 0, numAttrs)
columnIDs1 := make([]uint64, 0, numAttrs)
attrVals1 := make([]string, 0, numAttrs)
for n := 0; n < 1000000; n += 1000000 / numAttrs {
columnIDs0 = append(columnIDs0, uint64(n))
val0 := attrFun(uint64(n))
attrVals0 = append(attrVals0, val0)
setPql0 := fmt.Sprintf("Set(%d, %s=0) ", n, field)
if _, err := m0.API.Query(ctx, &pilosa.QueryRequest{Index: index, Query: setPql0}); err != nil {
t.Fatal(err)
}
columnIDs1 = append(columnIDs1, uint64(n+ShardWidth))
val1 := attrFun(uint64(n + ShardWidth))
attrVals1 = append(attrVals1, val1)
setPql1 := fmt.Sprintf("Set(%d, %s=0) ", n+ShardWidth, field)
if _, err := m1.API.Query(ctx, &pilosa.QueryRequest{Index: index, Query: setPql1}); err != nil {
t.Fatal(err)
}
}
// send shard0 to node1
req := &pilosa.ImportColumnAttrsRequest{
AttrKey: attrKey,
ColumnIDs: columnIDs0,
AttrVals: attrVals0,
Shard: 0,
Index: index,
}
if err := m1.API.ImportColumnAttrs(ctx, req); err != nil {
t.Fatal(err)
}
// send shard1 to node0
req = &pilosa.ImportColumnAttrsRequest{
AttrKey: attrKey,
ColumnIDs: columnIDs1,
AttrVals: attrVals1,
Shard: 1,
Index: index,
}
if err := m0.API.ImportColumnAttrs(ctx, req); err != nil {
t.Fatal(err)
}
// Query node0.
pql := fmt.Sprintf("Options(Row(%s=0), columnAttrs=true)", field)
res, err := m0.API.Query(ctx, &pilosa.QueryRequest{Index: index, Query: pql})
if err != nil {
t.Fatal(err)
}
if len(res.ColumnAttrSets) != 100 {
t.Fatal("incorrect number of column attrs set")
}
for _, v := range res.ColumnAttrSets {
attrVal := attrFun(v.ID)
if attrVal != v.Attrs[attrKey] {
t.Fatal(err)
}
}
// Query node1.
pql = fmt.Sprintf("Options(Row(%s=0), columnAttrs=true)", field)
res, err = m1.API.Query(ctx, &pilosa.QueryRequest{Index: index, Query: pql})
if err != nil {
t.Fatal(err)
}
if len(res.ColumnAttrSets) != 100 {
t.Fatal("incorrect number of column attrs set")
}
for _, v := range res.ColumnAttrSets {
attrVal := attrFun(v.ID)
if attrVal != v.Attrs[attrKey] {
t.Fatal(err)
}
}
})
}
func TestAPI_Import(t *testing.T) { func TestAPI_Import(t *testing.T) {
c := test.MustRunCluster(t, 2, c := test.MustRunCluster(t, 2,
[]server.CommandOption{ []server.CommandOption{
@ -175,10 +307,15 @@ func TestAPI_Import(t *testing.T) {
} }
// Query node1. // Query node1.
if err := test.RetryUntil(5*time.Second, func() error {
if res, err := m1.API.Query(ctx, &pilosa.QueryRequest{Index: index, Query: pql}); err != nil { if res, err := m1.API.Query(ctx, &pilosa.QueryRequest{Index: index, Query: pql}); err != nil {
t.Fatal(err) return err
} else if columns := res.Results[0].(*pilosa.Row).Columns(); !reflect.DeepEqual(columns, colIDs) { } else if columns := res.Results[0].(*pilosa.Row).Columns(); !reflect.DeepEqual(columns, colIDs) {
t.Fatalf("unexpected column ids: %+v", columns) return fmt.Errorf("unexpected column ids: %+v", columns)
}
return nil
}); err != nil {
t.Fatal(err)
} }
}) })
} }
@ -258,6 +395,202 @@ func TestAPI_ImportValue(t *testing.T) {
t.Fatal(err) t.Fatal(err)
} }
}) })
t.Run("ValDecimalField", func(t *testing.T) {
ctx := context.Background()
index := "valdec"
field := "fdec"
_, err := m1.API.CreateIndex(ctx, index, pilosa.IndexOptions{})
if err != nil {
t.Fatalf("creating index: %v", err)
}
fld, err := m1.API.CreateField(ctx, index, field, pilosa.OptFieldTypeDecimal(1))
if err != nil {
t.Fatalf("creating field: %v", err)
}
// Generate some keyed records.
values := []float64{}
colIDs := []uint64{}
for i := 0; i < 10; i++ {
values = append(values, float64(i)+0.1)
colIDs = append(colIDs, uint64(i))
}
// Import data with keys to the coordinator (node0) and verify that it gets
// translated and forwarded to the owner of shard 0 (node1; because of offsetModHasher)
req := &pilosa.ImportValueRequest{
Index: index,
Field: field,
ColumnIDs: colIDs,
FloatValues: values,
}
if err := m1.API.ImportValue(ctx, req); err != nil {
t.Fatal(err)
}
pql := fmt.Sprintf("Row(%s>6)", field)
// Query node0.
if res, err := m0.API.Query(ctx, &pilosa.QueryRequest{Index: index, Query: pql}); err != nil {
t.Fatal(err)
} else if ids := res.Results[0].(*pilosa.Row).Columns(); !reflect.DeepEqual(ids, colIDs[6:]) {
t.Fatalf("unexpected column keys: %+v", ids)
}
sum, count, err := fld.FloatSum(nil, field)
if err != nil {
t.Fatalf("getting floatsum: %v", err)
} else if sum != 0.1+1.1+2.1+3.1+4.1+5.1+6.1+7.1+8.1+9.1 {
t.Fatalf("unexpected sum: %f", sum)
} else if count != 10 {
t.Fatalf("unexpected count: %d", count)
}
min, count, err := fld.FloatMin(nil, field)
if err != nil {
t.Fatalf("getting floatmin: %v", err)
} else if min != 0.1 {
t.Fatalf("unexpected min: %f", min)
} else if count != 1 {
t.Fatalf("unexpected count: %d", count)
}
max, count, err := fld.FloatMax(nil, field)
if err != nil {
t.Fatalf("getting floatmax: %v", err)
} else if max != 9.1 {
t.Fatalf("unexpected max: %f", max)
} else if count != 1 {
t.Fatalf("unexpected count: %d", count)
}
val, exists, err := fld.FloatValue(1)
if err != nil {
t.Fatalf("unepxected err getting floatvalue")
} else if !exists {
t.Fatalf("column 1 should exist")
} else if val != 1.1 {
t.Fatalf("unexpected floatvalue %f", val)
}
changed, err := fld.SetFloatValue(11, 11.1)
if err != nil {
t.Fatalf("setting float value: %v", err)
} else if !changed {
t.Fatalf("expected change")
}
val, exists, err = fld.FloatValue(11)
if err != nil {
t.Fatalf("getting float val: %v", err)
} else if !exists {
t.Fatalf("should exist")
} else if val != 11.1 {
t.Fatalf("unexpected val: %f", 11.1)
}
})
t.Run("ValDecimalFieldNegativeScale", func(t *testing.T) {
ctx := context.Background()
index := "valdecneg"
field := "fdecneg"
_, err := m0.API.CreateIndex(ctx, index, pilosa.IndexOptions{})
if err != nil {
t.Fatalf("creating index: %v", err)
}
_, err = m0.API.CreateField(ctx, index, field, pilosa.OptFieldTypeDecimal(-1))
if err != nil {
t.Fatalf("creating field: %v", err)
}
// Generate some keyed records.
values := []float64{}
colIDs := []uint64{}
for i := 0; i < 10; i++ {
values = append(values, float64(i)*100+10)
colIDs = append(colIDs, uint64(i))
}
// Import data with keys to the coordinator (node0) and verify that it gets
// translated and forwarded to the owner of shard 0 (node1; because of offsetModHasher)
req := &pilosa.ImportValueRequest{
Index: index,
Field: field,
ColumnIDs: colIDs,
FloatValues: values,
}
if err := m1.API.ImportValue(ctx, req); err != nil {
t.Fatal(err)
}
pql := fmt.Sprintf("Row(%s>600)", field)
// Query node0.
if res, err := m0.API.Query(ctx, &pilosa.QueryRequest{Index: index, Query: pql}); err != nil {
t.Fatal(err)
} else if ids := res.Results[0].(*pilosa.Row).Columns(); !reflect.DeepEqual(ids, colIDs[6:]) {
t.Fatalf("unexpected column keys: %+v", ids)
}
})
t.Run("ValStringField", func(t *testing.T) {
ctx := context.Background()
index := "valstr"
field := "fstr"
fgnIndex := "fgnvalstr"
_, err := m0.API.CreateIndex(ctx, index, pilosa.IndexOptions{})
if err != nil {
t.Fatalf("creating index: %v", err)
}
_, err = m0.API.CreateIndex(ctx, fgnIndex, pilosa.IndexOptions{Keys: true})
if err != nil {
t.Fatalf("creating foreign index: %v", err)
}
_, err = m0.API.CreateField(ctx, index, field,
pilosa.OptFieldTypeInt(0, math.MaxInt64),
pilosa.OptFieldForeignIndex(fgnIndex),
)
if err != nil {
t.Fatalf("creating field: %v", err)
}
// Generate some keyed records.
values := []string{}
colIDs := []uint64{}
for i := 0; i < 10; i++ {
value := fmt.Sprintf("strval-%d", (i)*100+10)
values = append(values, value)
colIDs = append(colIDs, uint64(i))
}
// Import data with keys to the coordinator (node0) and verify that it gets
// translated and forwarded to the owner of shard 0 (node1; because of offsetModHasher)
req := &pilosa.ImportValueRequest{
Index: index,
Field: field,
ColumnIDs: colIDs,
StringValues: values,
}
if err := m0.API.ImportValue(ctx, req); err != nil {
t.Fatal(err)
}
pql := fmt.Sprintf(`Row(%s=="strval-110")`, field)
// Query node0.
if res, err := m0.API.Query(ctx, &pilosa.QueryRequest{Index: index, Query: pql}); err != nil {
t.Fatal(err)
} else if ids := res.Results[0].(*pilosa.Row).Columns(); !reflect.DeepEqual(ids, []uint64{1}) {
t.Fatalf("unexpected columns: %+v", ids)
}
})
} }
// offsetModHasher represents a simple, mod-based hashing offset by 1. // offsetModHasher represents a simple, mod-based hashing offset by 1.

View file

@ -16,6 +16,7 @@ package pilosa
import ( import (
"bytes" "bytes"
"encoding/json"
"fmt" "fmt"
"io" "io"
"sort" "sort"
@ -172,6 +173,7 @@ func (c *rankCache) Add(id uint64, n uint64) {
// unless the count is 0, which is effectively used // unless the count is 0, which is effectively used
// to clear the cache value. // to clear the cache value.
if n < c.thresholdValue && n > 0 { if n < c.thresholdValue && n > 0 {
delete(c.entries, id)
return return
} }
@ -185,6 +187,7 @@ func (c *rankCache) BulkAdd(id uint64, n uint64) {
c.mu.Lock() c.mu.Lock()
defer c.mu.Unlock() defer c.mu.Unlock()
if n < c.thresholdValue { if n < c.thresholdValue {
delete(c.entries, id)
return return
} }
@ -320,6 +323,18 @@ type Pair struct {
Count uint64 `json:"count"` Count uint64 `json:"count"`
} }
// PairField
type PairField struct {
Pair Pair
Field string
}
// MarshalJSON marshals PairField into a JSON-encoded byte slice,
// excluding `Field`.
func (p PairField) MarshalJSON() ([]byte, error) {
return json.Marshal(p.Pair)
}
// Pairs is a sortable slice of Pair objects. // Pairs is a sortable slice of Pair objects.
type Pairs []Pair type Pairs []Pair
@ -395,6 +410,18 @@ func (p Pairs) String() string {
return buf.String() return buf.String()
} }
// PairsField
type PairsField struct {
Pairs []Pair
Field string
}
// MarshalJSON marshals PairsField into a JSON-encoded byte slice,
// excluding `Field`.
func (p PairsField) MarshalJSON() ([]byte, error) {
return json.Marshal(p.Pairs)
}
// uint64Slice represents a sortable slice of uint64 numbers. // uint64Slice represents a sortable slice of uint64 numbers.
type uint64Slice []uint64 type uint64Slice []uint64

View file

@ -20,8 +20,8 @@ import (
"github.com/pilosa/pilosa/v2" "github.com/pilosa/pilosa/v2"
) )
// Ensure a bitmap query can be executed. // Ensure cache stays constrained to its configured size.
func TestCache_Rank(t *testing.T) { func TestCache_Rank_Size(t *testing.T) {
cacheSize := uint32(3) cacheSize := uint32(3)
cache := pilosa.NewRankCache(cacheSize) cache := pilosa.NewRankCache(cacheSize)
for i := 1; i < int(2*cacheSize); i++ { for i := 1; i < int(2*cacheSize); i++ {
@ -31,5 +31,26 @@ func TestCache_Rank(t *testing.T) {
if cache.Len() != int(cacheSize) { if cache.Len() != int(cacheSize) {
t.Fatalf("unexpected cache Size: %d!=%d expected\n", cache.Len(), cacheSize) t.Fatalf("unexpected cache Size: %d!=%d expected\n", cache.Len(), cacheSize)
} }
}
// Ensure cache entries set below threshold are handled appropriately.
func TestCache_Rank_Threshold(t *testing.T) {
cacheSize := uint32(5)
cache := pilosa.NewRankCache(cacheSize)
for i := 1; i < int(2*cacheSize); i++ {
cache.Add(uint64(i), 3)
}
// Set the cache value for rows 4 and 5 to a number below the threshold
// value (which is 3), and ensure that they gets zeroed out.
cache.Add(4, 1)
cache.BulkAdd(5, 1)
cache.Recalculate()
if cache.Get(4) != 0 {
t.Fatalf("unexpected cache value after Add: %d!=%d expected\n", cache.Get(4), 0)
}
if cache.Get(5) != 0 {
t.Fatalf("unexpected cache value after BulkAdd: %d!=%d expected\n", cache.Get(5), 0)
}
} }

View file

@ -59,6 +59,7 @@ type InternalClient interface {
EnsureFieldWithOptions(ctx context.Context, index, field string, opt FieldOptions) error EnsureFieldWithOptions(ctx context.Context, index, field string, opt FieldOptions) error
ImportValue(ctx context.Context, index, field string, shard uint64, vals []FieldValue, opts ...ImportOption) error ImportValue(ctx context.Context, index, field string, shard uint64, vals []FieldValue, opts ...ImportOption) error
ImportValueK(ctx context.Context, index, field string, vals []FieldValue, opts ...ImportOption) error ImportValueK(ctx context.Context, index, field string, vals []FieldValue, opts ...ImportOption) error
ImportValue2(ctx context.Context, req *ImportValueRequest, options *ImportOptions) error
ExportCSV(ctx context.Context, index, field string, shard uint64, w io.Writer) error ExportCSV(ctx context.Context, index, field string, shard uint64, w io.Writer) error
CreateField(ctx context.Context, index, field string) error CreateField(ctx context.Context, index, field string) error
CreateFieldWithOptions(ctx context.Context, index, field string, opt FieldOptions) error CreateFieldWithOptions(ctx context.Context, index, field string, opt FieldOptions) error
@ -69,6 +70,7 @@ type InternalClient interface {
SendMessage(ctx context.Context, uri *URI, msg []byte) error SendMessage(ctx context.Context, uri *URI, msg []byte) error
RetrieveShardFromURI(ctx context.Context, index, field, view string, shard uint64, uri URI) (io.ReadCloser, error) RetrieveShardFromURI(ctx context.Context, index, field, view string, shard uint64, uri URI) (io.ReadCloser, error)
ImportRoaring(ctx context.Context, uri *URI, index, field string, shard uint64, remote bool, req *ImportRoaringRequest) error ImportRoaring(ctx context.Context, uri *URI, index, field string, shard uint64, remote bool, req *ImportRoaringRequest) error
ImportColumnAttrs(ctx context.Context, uri *URI, index string, req *ImportColumnAttrsRequest) error
} }
//=============== //===============
@ -129,9 +131,18 @@ func (n nopInternalClient) Import(ctx context.Context, index, field string, shar
func (n nopInternalClient) ImportK(ctx context.Context, index, field string, bits []Bit, opts ...ImportOption) error { func (n nopInternalClient) ImportK(ctx context.Context, index, field string, bits []Bit, opts ...ImportOption) error {
return nil return nil
} }
func (n nopInternalClient) ImportValue2(ctx context.Context, req *ImportValueRequest, options *ImportOptions) error {
return nil
}
func (n nopInternalClient) ImportRoaring(ctx context.Context, uri *URI, index, field string, shard uint64, remote bool, req *ImportRoaringRequest) error { func (n nopInternalClient) ImportRoaring(ctx context.Context, uri *URI, index, field string, shard uint64, remote bool, req *ImportRoaringRequest) error {
return nil return nil
} }
func (n nopInternalClient) ImportColumnAttrs(ctx context.Context, uri *URI, index string, req *ImportColumnAttrsRequest) error {
return nil
}
func (n nopInternalClient) EnsureIndex(ctx context.Context, name string, options IndexOptions) error { func (n nopInternalClient) EnsureIndex(ctx context.Context, name string, options IndexOptions) error {
return nil return nil
} }

View file

@ -158,6 +158,7 @@ func TestFragSources(t *testing.T) {
c5.addNodeBasicSorted(node3) c5.addNodeBasicSorted(node3)
idx := newIndexWithTempPath("i") idx := newIndexWithTempPath("i")
defer idx.Close()
field, err := idx.CreateFieldIfNotExists("f", OptFieldTypeDefault()) field, err := idx.CreateFieldIfNotExists("f", OptFieldTypeDefault())
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
@ -948,6 +949,9 @@ func TestCluster_confirmNodeDownUp(t *testing.T) {
} }
func TestCluster_confirmNodeDownTimeout(t *testing.T) { func TestCluster_confirmNodeDownTimeout(t *testing.T) {
if testing.Short() {
t.Skip()
}
r := mux.NewRouter() r := mux.NewRouter()
r.HandleFunc("/version", http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { r.HandleFunc("/version", http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
time.Sleep(confirmDownSleep * time.Second * confirmDownRetries) time.Sleep(confirmDownSleep * time.Second * confirmDownRetries)
@ -973,10 +977,12 @@ func TestCluster_confirmNodeDownTimeout(t *testing.T) {
if !confirmNodeDown(uri, logger.NewVerboseLogger(os.Stdout)) { if !confirmNodeDown(uri, logger.NewVerboseLogger(os.Stdout)) {
t.Errorf("expected node to be down") t.Errorf("expected node to be down")
} }
} }
func TestCluster_confirmNodeDownDown(t *testing.T) { func TestCluster_confirmNodeDownDown(t *testing.T) {
if testing.Short() {
t.Skip()
}
uri := URI{} uri := URI{}
uri.Scheme = "http" uri.Scheme = "http"
uri.Host = "DoesntMatter" uri.Host = "DoesntMatter"
@ -985,5 +991,4 @@ func TestCluster_confirmNodeDownDown(t *testing.T) {
if !confirmNodeDown(uri, logger.NewVerboseLogger(os.Stdout)) { if !confirmNodeDown(uri, logger.NewVerboseLogger(os.Stdout)) {
t.Errorf("expected node to be down") t.Errorf("expected node to be down")
} }
} }

View file

@ -53,7 +53,7 @@ omitted. If it is present then its format should be YYYY-MM-DDTHH:MM.
flags.StringVarP(&Importer.Field, "field", "f", "", "Field to import into.") flags.StringVarP(&Importer.Field, "field", "f", "", "Field to import into.")
flags.BoolVar(&Importer.IndexOptions.Keys, "index-keys", false, "Specify keys=true when creating an index") flags.BoolVar(&Importer.IndexOptions.Keys, "index-keys", false, "Specify keys=true when creating an index")
flags.BoolVar(&Importer.FieldOptions.Keys, "field-keys", false, "Specify keys=true when creating a field") flags.BoolVar(&Importer.FieldOptions.Keys, "field-keys", false, "Specify keys=true when creating a field")
flags.StringVar(&Importer.FieldOptions.Type, "field-type", "", "Specify the field type when creating a field. One of: set, int, time, bool, mutex") flags.StringVar(&Importer.FieldOptions.Type, "field-type", "", "Specify the field type when creating a field. One of: set, int, decimal, time, bool, mutex")
flags.Int64Var(&Importer.FieldOptions.Min, "field-min", 0, "Specify the minimum for an int field on creation") flags.Int64Var(&Importer.FieldOptions.Min, "field-min", 0, "Specify the minimum for an int field on creation")
flags.Int64Var(&Importer.FieldOptions.Max, "field-max", 0, "Specify the maximum for an int field on creation") flags.Int64Var(&Importer.FieldOptions.Max, "field-max", 0, "Specify the maximum for an int field on creation")
flags.StringVar(&Importer.FieldOptions.CacheType, "field-cache-type", pilosa.CacheTypeRanked, "Specify the cache type for a set field on creation. One of: none, lru, ranked") flags.StringVar(&Importer.FieldOptions.CacheType, "field-cache-type", pilosa.CacheTypeRanked, "Specify the cache type for a set field on creation. One of: none, lru, ranked")

View file

@ -42,7 +42,7 @@ func TestServerConfig(t *testing.T) {
tests := []commandTest{ tests := []commandTest{
// TEST 0 // TEST 0
{ {
args: []string{"server", "--data-dir", actualDataDir, "--cluster.hosts", "localhost:42454,localhost:10110", "--bind", "localhost:42454", "--translation.map-size", "100000"}, args: []string{"server", "--data-dir", actualDataDir, "--cluster.hosts", "localhost:42454,localhost:10110", "--bind", "localhost:42454", "--bind-grpc", "localhost:30112", "--translation.map-size", "100000"},
env: map[string]string{ env: map[string]string{
"PILOSA_DATA_DIR": "/tmp/myEnvDatadir", "PILOSA_DATA_DIR": "/tmp/myEnvDatadir",
"PILOSA_CLUSTER_LONG_QUERY_TIME": "1m30s", "PILOSA_CLUSTER_LONG_QUERY_TIME": "1m30s",
@ -53,6 +53,7 @@ func TestServerConfig(t *testing.T) {
cfgFileContent: ` cfgFileContent: `
data-dir = "/tmp/myFileDatadir" data-dir = "/tmp/myFileDatadir"
bind = "localhost:0" bind = "localhost:0"
bind-grpc = "localhost:0"
max-writes-per-request = 3000 max-writes-per-request = 3000
[cluster] [cluster]
@ -96,6 +97,7 @@ func TestServerConfig(t *testing.T) {
}, },
cfgFileContent: ` cfgFileContent: `
bind = "localhost:0" bind = "localhost:0"
bind-grpc = "localhost:0"
data-dir = "` + actualDataDir + `" data-dir = "` + actualDataDir + `"
[cluster] [cluster]
disabled = true disabled = true
@ -122,6 +124,7 @@ func TestServerConfig(t *testing.T) {
env: map[string]string{}, env: map[string]string{},
cfgFileContent: ` cfgFileContent: `
bind = "localhost:19444" bind = "localhost:19444"
bind-grpc = "localhost:29444"
data-dir = "` + actualDataDir + `" data-dir = "` + actualDataDir + `"
[cluster] [cluster]
hosts = [ hosts = [

View file

@ -20,6 +20,7 @@ import (
"fmt" "fmt"
"io" "io"
"log" "log"
"math"
"os" "os"
"sort" "sort"
"strconv" "strconv"
@ -102,11 +103,13 @@ func (cmd *ImportCommand) Run(ctx context.Context) error {
if cmd.FieldOptions.Type == "" { if cmd.FieldOptions.Type == "" {
// set the correct type for the field // set the correct type for the field
if cmd.FieldOptions.TimeQuantum != "" { if cmd.FieldOptions.TimeQuantum != "" {
cmd.FieldOptions.Type = "time" cmd.FieldOptions.Type = pilosa.FieldTypeTime
} else if cmd.FieldOptions.Min != 0 || cmd.FieldOptions.Max != 0 { } else if cmd.FieldOptions.Min != 0 || cmd.FieldOptions.Max != 0 {
cmd.FieldOptions.Type = "int" cmd.FieldOptions.Type = pilosa.FieldTypeInt
} else { } else {
cmd.FieldOptions.Type = "set" cmd.FieldOptions.Type = pilosa.FieldTypeSet
cmd.FieldOptions.CacheType = pilosa.CacheTypeRanked
cmd.FieldOptions.CacheSize = pilosa.DefaultCacheSize
} }
} }
err := cmd.ensureSchema(ctx) err := cmd.ensureSchema(ctx)
@ -163,8 +166,8 @@ func (cmd *ImportCommand) ensureSchema(ctx context.Context) error {
// importPath parses a path into bits and imports it to the server. // importPath parses a path into bits and imports it to the server.
func (cmd *ImportCommand) importPath(ctx context.Context, fieldType string, useColumnKeys, useRowKeys bool, path string) error { func (cmd *ImportCommand) importPath(ctx context.Context, fieldType string, useColumnKeys, useRowKeys bool, path string) error {
// If fieldType is `int`, treat the import data as values to be range-encoded. // If fieldType is `int`, treat the import data as values to be range-encoded.
if fieldType == pilosa.FieldTypeInt { if fieldType == pilosa.FieldTypeInt || fieldType == pilosa.FieldTypeDecimal {
return cmd.bufferValues(ctx, useColumnKeys, path) return cmd.bufferValues(ctx, useColumnKeys, fieldType == pilosa.FieldTypeDecimal, path)
} }
return cmd.bufferBits(ctx, useColumnKeys, useRowKeys, path) return cmd.bufferBits(ctx, useColumnKeys, useRowKeys, path)
} }
@ -285,9 +288,13 @@ func (cmd *ImportCommand) importBits(ctx context.Context, useColumnKeys, useRowK
return nil return nil
} }
// bufferValues buffers slices of FieldValues to be imported as a batch. // bufferValues buffers slices of record identifiers and values to be imported as a batch.
func (cmd *ImportCommand) bufferValues(ctx context.Context, useColumnKeys bool, path string) error { func (cmd *ImportCommand) bufferValues(ctx context.Context, useColumnKeys, parseAsFloat bool, path string) error {
a := make([]pilosa.FieldValue, 0, cmd.BufferSize) req := &pilosa.ImportValueRequest{
Index: cmd.Index,
Field: cmd.Field,
Shard: math.MaxUint64,
}
var r *csv.Reader var r *csv.Reader
@ -307,6 +314,7 @@ func (cmd *ImportCommand) bufferValues(ctx context.Context, useColumnKeys bool,
r.FieldsPerRecord = -1 r.FieldsPerRecord = -1
rnum := 0 rnum := 0
for { for {
rnum++ rnum++
@ -325,69 +333,44 @@ func (cmd *ImportCommand) bufferValues(ctx context.Context, useColumnKeys bool,
return fmt.Errorf("bad column count on row %d: col=%d", rnum, len(record)) return fmt.Errorf("bad column count on row %d: col=%d", rnum, len(record))
} }
var val pilosa.FieldValue
// Parse column id. // Parse column id.
if useColumnKeys { if useColumnKeys {
val.ColumnKey = record[0] req.ColumnKeys = append(req.ColumnKeys, record[0])
} else if columnID, err := strconv.ParseUint(record[0], 10, 64); err == nil {
req.ColumnIDs = append(req.ColumnIDs, columnID)
} else { } else {
if val.ColumnID, err = strconv.ParseUint(record[0], 10, 64); err != nil {
return fmt.Errorf("invalid column id on row %d: %q", rnum, record[0]) return fmt.Errorf("invalid column id on row %d: %q", rnum, record[0])
} }
}
// Parse FieldValue. // Parse value.
if parseAsFloat {
value, err := strconv.ParseFloat(record[1], 64)
if err != nil {
return errors.Wrapf(err, "parseing value '%s' as float", record[1])
}
req.FloatValues = append(req.FloatValues, value)
} else {
value, err := strconv.ParseInt(record[1], 10, 64) value, err := strconv.ParseInt(record[1], 10, 64)
if err != nil { if err != nil {
return fmt.Errorf("invalid value on row %d: %q", rnum, record[1]) return errors.Wrapf(err, "invalid value on row %d: %q", rnum, record[1])
} }
val.Value = value req.Values = append(req.Values, value)
a = append(a, val)
// If we've reached the buffer size then import FieldValues.
if len(a) == cmd.BufferSize {
if err := cmd.importValues(ctx, useColumnKeys, a); err != nil {
return err
} }
a = a[:0]
// If we've reached the buffer size then import the batch.
if len(req.ColumnKeys) == cmd.BufferSize || len(req.ColumnIDs) == cmd.BufferSize {
if err := cmd.client.ImportValue2(ctx, req, &pilosa.ImportOptions{}); err != nil {
return errors.Wrap(err, "importing values")
}
req.ColumnIDs = req.ColumnIDs[:0]
req.ColumnKeys = req.ColumnKeys[:0]
req.Values = req.Values[:0]
req.FloatValues = req.FloatValues[:0]
} }
} }
// If there are still values in the buffer then flush them. // If there are still values in the buffer then flush them.
return cmd.importValues(ctx, useColumnKeys, a) return errors.Wrap(cmd.client.ImportValue2(ctx, req, &pilosa.ImportOptions{}), "importing values")
}
// importValues sends batches of FieldValues to the server.
func (cmd *ImportCommand) importValues(ctx context.Context, useColumnKeys bool, vals []pilosa.FieldValue) error {
logger := log.New(cmd.Stderr, "", log.LstdFlags)
// If keys are used, all values are sent to the primary translate store (i.e. coordinator).
if useColumnKeys {
logger.Printf("importing keyed values: n=%d", len(vals))
if err := cmd.client.ImportValueK(ctx, cmd.Index, cmd.Field, vals); err != nil {
return errors.Wrap(err, "importing keys")
}
return nil
}
// Group vals by shard.
logger.Printf("grouping %d vals", len(vals))
valsByShard := http.FieldValues(vals).GroupByShard()
// Parse path into FieldValues.
for shard, vals := range valsByShard {
if cmd.Sort {
sort.Sort(http.FieldValues(vals))
}
logger.Printf("importing shard: %d, n=%d", shard, len(vals))
if err := cmd.client.ImportValue(ctx, cmd.Index, cmd.Field, shard, vals, pilosa.OptImportOptionsClear(cmd.Clear)); err != nil {
return errors.Wrap(err, "importing values")
}
}
return nil
} }
func (cmd *ImportCommand) TLSHost() string { func (cmd *ImportCommand) TLSHost() string {

View file

@ -26,6 +26,7 @@ func BuildServerFlags(cmd *cobra.Command, srv *server.Command) {
flags := cmd.Flags() flags := cmd.Flags()
flags.StringVarP(&srv.Config.DataDir, "data-dir", "d", srv.Config.DataDir, "Directory to store pilosa data files.") flags.StringVarP(&srv.Config.DataDir, "data-dir", "d", srv.Config.DataDir, "Directory to store pilosa data files.")
flags.StringVarP(&srv.Config.Bind, "bind", "b", srv.Config.Bind, "Default URI on which pilosa should listen.") flags.StringVarP(&srv.Config.Bind, "bind", "b", srv.Config.Bind, "Default URI on which pilosa should listen.")
flags.StringVar(&srv.Config.BindGRPC, "bind-grpc", srv.Config.BindGRPC, "URI on which pilosa should listen for gRPC requests.")
flags.StringVar(&srv.Config.Advertise, "advertise", srv.Config.Advertise, "Address to advertise externally.") flags.StringVar(&srv.Config.Advertise, "advertise", srv.Config.Advertise, "Address to advertise externally.")
flags.IntVarP(&srv.Config.MaxWritesPerRequest, "max-writes-per-request", "", srv.Config.MaxWritesPerRequest, "Number of write commands per request.") flags.IntVarP(&srv.Config.MaxWritesPerRequest, "max-writes-per-request", "", srv.Config.MaxWritesPerRequest, "Number of write commands per request.")
flags.StringVar(&srv.Config.LogPath, "log-path", srv.Config.LogPath, "Log path") flags.StringVar(&srv.Config.LogPath, "log-path", srv.Config.LogPath, "Log path")

View file

@ -233,7 +233,7 @@ func (d *diagnosticsCollector) EnrichWithSchemaProperties() {
numIndexes++ numIndexes++
for _, field := range index.Fields() { for _, field := range index.Fields() {
numFields++ numFields++
if field.Type() == FieldTypeInt { if field.Type() == FieldTypeInt || field.Type() == FieldTypeDecimal {
bsiFieldCount++ bsiFieldCount++
} }
if field.TimeQuantum() != "" { if field.TimeQuantum() != "" {

View file

@ -100,7 +100,7 @@ curl localhost:10101/index/repository -X POST
``` response ``` response
{"success":true} {"success":true}
``` ```
The index name must be 64 characters or fewer, start with a letter, and consist only of lowercase alphanumeric characters or `_-`. The same goes for field names. The index name must be 230 characters or fewer, start with a letter, and consist only of lowercase alphanumeric characters or `_-`. The same goes for field names.
Let's create the `stargazer` field which has user IDs of stargazers as its rows: Let's create the `stargazer` field which has user IDs of stargazers as its rows:
``` request ``` request
@ -325,7 +325,7 @@ Next, let's create the `repository` index:
repository := schema.Index("repository") repository := schema.Index("repository")
``` ```
The index name must be 64 characters or fewer, start with a letter, and consist only of lowercase alphanumeric characters or `_-`. The same goes for field names. The index name must be 230 characters or fewer, start with a letter, and consist only of lowercase alphanumeric characters or `_-`. The same goes for field names.
Let's create the `stargazer` field which has user IDs of stargazers as its rows: Let's create the `stargazer` field which has user IDs of stargazers as its rows:
``` ```
@ -615,7 +615,7 @@ Next, let's create the `repository` index:
``` ```
Index repository = schema.index("repository"); Index repository = schema.index("repository");
``` ```
The index name must be 64 characters or fewer, start with a letter, and consist only of lowercase alphanumeric characters or `_-`. The same goes for field names. The index name must be 230 characters or fewer, start with a letter, and consist only of lowercase alphanumeric characters or `_-`. The same goes for field names.
Let's create the `stargazer` field which has user IDs of stargazers as its rows: Let's create the `stargazer` field which has user IDs of stargazers as its rows:
``` ```
@ -818,7 +818,7 @@ Next, let's create the `repository` index:
``` ```
repository = schema.index("repository") repository = schema.index("repository")
``` ```
The index name must be 64 characters or fewer, start with a letter, and consist only of lowercase alphanumeric characters or `_-`. The same goes for field names. The index name must be 230 characters or fewer, start with a letter, and consist only of lowercase alphanumeric characters or `_-`. The same goes for field names.
Let's create the `stargazer` field which has user IDs of stargazers as its rows: Let's create the `stargazer` field which has user IDs of stargazers as its rows:
``` ```

View file

@ -43,7 +43,7 @@ curl localhost:10101/index/repository/query \
#### Arguments and Types #### Arguments and Types
* `field` The field specifies on which Pilosa [field](../glossary/#field) the query will operate. Valid field names are lower case strings; they start with a lowercase letter, and contain only alphanumeric characters and `_-`. They must be 64 characters or less in length. * `field` The field specifies on which Pilosa [field](../glossary/#field) the query will operate. Valid field names are lower case strings; they start with a lowercase letter, and contain only alphanumeric characters and `_-`. They must be 230 characters or less in length.
* `TIMESTAMP` This is a timestamp in the following format `YYYY-MM-DDTHH:MM` (e.g. 2006-01-02T15:04). * `TIMESTAMP` This is a timestamp in the following format `YYYY-MM-DDTHH:MM` (e.g. 2006-01-02T15:04).
* `UINT` An unsigned integer (e.g. 42839). * `UINT` An unsigned integer (e.g. 42839).
* `BOOL` A boolean value, `true` or `false`. * `BOOL` A boolean value, `true` or `false`.
@ -864,7 +864,7 @@ Rows(job)
**Spec:** **Spec:**
``` ```
GroupBy(<ROWS_CALL>, [<ROWS_CALL>...], limit=<UINT>, filter=<ROW_CALL>) GroupBy(<ROWS_CALL>, [<ROWS_CALL>...], limit=<UINT>, filter=<ROW_CALL>, aggregate=<CALL>)
``` ```
**Description:** **Description:**
@ -874,14 +874,18 @@ taking one row each from the specified `Rows` calls. It returns only those
combinations for which the count is greater than 0. combinations for which the count is greater than 0.
The optional `filter` argument takes any type of `Row` query (e.g. Row, Union, The optional `filter` argument takes any type of `Row` query (e.g. Row, Union,
Intersect, etc.) which will be intersected with each result prior to returning Intersect, etc.) which will be intersected with each result prior to returning
the count. This is analogous to a WHERE clause applied to a relational GROUP BY the count. This is analagous to a WHERE clause applied to a relational GROUP BY
query. query.
The optional `limit` argument limits the number of results returned. The results The optional `limit` argument limits the number of results returned. The results
are ordered, so as long as the data isn't changing, the same query will return are ordered, so as long as the data isn't changing, the same query will return
the same result set. the same result set.
The optional `aggregate` argument takes a `Sum()` query which will be used to
calculate the sum & count of each group. This is similar to using a `SUM()` in
the SELECT clause of a relation GROUP BY query.
Paging through results is supported by passing the `previous` argument to each Paging through results is supported by passing the `previous` argument to each
of the `Rows` calls in the GroupBy. Take the last result from your previous of the `Rows` calls in the GroupBy. Take the last result from your previous
`GroupBy` query, and pass each row ID in that result as the `previous` argument `GroupBy` query, and pass each row ID in that result as the `previous` argument

View file

@ -225,6 +225,14 @@ func (Serializer) Unmarshal(buf []byte, m pilosa.Message) error {
} }
decodeImportRoaringRequest(msg, mt) decodeImportRoaringRequest(msg, mt)
return nil return nil
case *pilosa.ImportColumnAttrsRequest:
msg := &internal.ImportColumnAttrsRequest{}
err := proto.Unmarshal(buf, msg)
if err != nil {
return errors.Wrap(err, "unmarshaling ImportColumnAttrsRequest")
}
decodeImportColumnAttrsRequest(msg, mt)
return nil
case *pilosa.ImportResponse: case *pilosa.ImportResponse:
msg := &internal.ImportResponse{} msg := &internal.ImportResponse{}
err := proto.Unmarshal(buf, msg) err := proto.Unmarshal(buf, msg)
@ -318,6 +326,8 @@ func encodeToProto(m pilosa.Message) proto.Message {
return encodeImportValueRequest(mt) return encodeImportValueRequest(mt)
case *pilosa.ImportRoaringRequest: case *pilosa.ImportRoaringRequest:
return encodeImportRoaringRequest(mt) return encodeImportRoaringRequest(mt)
case *pilosa.ImportColumnAttrsRequest:
return encodeImportColumnAttrsRequest(mt)
case *pilosa.ImportResponse: case *pilosa.ImportResponse:
return encodeImportResponse(mt) return encodeImportResponse(mt)
case *pilosa.BlockDataRequest: case *pilosa.BlockDataRequest:
@ -375,6 +385,8 @@ func encodeImportValueRequest(m *pilosa.ImportValueRequest) *internal.ImportValu
ColumnIDs: m.ColumnIDs, ColumnIDs: m.ColumnIDs,
ColumnKeys: m.ColumnKeys, ColumnKeys: m.ColumnKeys,
Values: m.Values, Values: m.Values,
FloatValues: m.FloatValues,
StringValues: m.StringValues,
} }
} }
@ -394,15 +406,30 @@ func encodeImportRoaringRequest(m *pilosa.ImportRoaringRequest) *internal.Import
} }
} }
func encodeImportColumnAttrsRequest(m *pilosa.ImportColumnAttrsRequest) *internal.ImportColumnAttrsRequest {
return &internal.ImportColumnAttrsRequest{
Index: m.Index,
Shard: m.Shard,
AttrKey: m.AttrKey,
AttrVals: m.AttrVals,
ColumnIDs: m.ColumnIDs,
}
}
func encodeQueryRequest(m *pilosa.QueryRequest) *internal.QueryRequest { func encodeQueryRequest(m *pilosa.QueryRequest) *internal.QueryRequest {
return &internal.QueryRequest{ r := &internal.QueryRequest{
Query: m.Query, Query: m.Query,
Shards: m.Shards, Shards: m.Shards,
ColumnAttrs: m.ColumnAttrs, ColumnAttrs: m.ColumnAttrs,
Remote: m.Remote, Remote: m.Remote,
ExcludeRowAttrs: m.ExcludeRowAttrs, ExcludeRowAttrs: m.ExcludeRowAttrs,
ExcludeColumns: m.ExcludeColumns, ExcludeColumns: m.ExcludeColumns,
EmbeddedData: make([]*internal.Row, len(m.EmbeddedData)),
} }
for i := range m.EmbeddedData {
r.EmbeddedData[i] = encodeRow(m.EmbeddedData[i])
}
return r
} }
func encodeQueryResponse(m *pilosa.QueryResponse) *internal.QueryResponse { func encodeQueryResponse(m *pilosa.QueryResponse) *internal.QueryResponse {
@ -415,12 +442,18 @@ func encodeQueryResponse(m *pilosa.QueryResponse) *internal.QueryResponse {
pb.Results[i] = &internal.QueryResult{} pb.Results[i] = &internal.QueryResult{}
switch result := m.Results[i].(type) { switch result := m.Results[i].(type) {
case pilosa.SignedRow:
pb.Results[i].Type = queryResultTypeSignedRow
pb.Results[i].SignedRow = encodeSignedRow(result)
case *pilosa.Row: case *pilosa.Row:
pb.Results[i].Type = queryResultTypeRow pb.Results[i].Type = queryResultTypeRow
pb.Results[i].Row = encodeRow(result) pb.Results[i].Row = encodeRow(result)
case []pilosa.Pair: case []pilosa.Pair:
pb.Results[i].Type = queryResultTypePairs pb.Results[i].Type = queryResultTypePairs
pb.Results[i].Pairs = encodePairs(result) pb.Results[i].Pairs = encodePairs(result)
case *pilosa.PairsField:
pb.Results[i].Type = queryResultTypePairsField
pb.Results[i].PairsField = encodePairsField(result)
case pilosa.ValCount: case pilosa.ValCount:
pb.Results[i].Type = queryResultTypeValCount pb.Results[i].Type = queryResultTypeValCount
pb.Results[i].ValCount = encodeValCount(result) pb.Results[i].ValCount = encodeValCount(result)
@ -442,10 +475,13 @@ func encodeQueryResponse(m *pilosa.QueryResponse) *internal.QueryResponse {
case pilosa.Pair: case pilosa.Pair:
pb.Results[i].Type = queryResultTypePair pb.Results[i].Type = queryResultTypePair
pb.Results[i].Pairs = []*internal.Pair{encodePair(result)} pb.Results[i].Pairs = []*internal.Pair{encodePair(result)}
case pilosa.PairField:
pb.Results[i].Type = queryResultTypePairField
pb.Results[i].Pairs = []*internal.Pair{encodePairField(result)}
case nil: case nil:
pb.Results[i].Type = queryResultTypeNil pb.Results[i].Type = queryResultTypeNil
default: default:
panic(fmt.Errorf("unknown type: %d", pb.Results[i].Type)) panic(fmt.Errorf("unknown type: %T", m.Results[i]))
} }
} }
@ -538,9 +574,11 @@ func encodeFieldOptions(o *pilosa.FieldOptions) *internal.FieldOptions {
Min: o.Min, Min: o.Min,
Max: o.Max, Max: o.Max,
Base: o.Base, Base: o.Base,
Scale: o.Scale,
BitDepth: uint64(o.BitDepth), BitDepth: uint64(o.BitDepth),
TimeQuantum: string(o.TimeQuantum), TimeQuantum: string(o.TimeQuantum),
Keys: o.Keys, Keys: o.Keys,
ForeignIndex: o.ForeignIndex,
} }
} }
@ -808,9 +846,11 @@ func decodeFieldOptions(options *internal.FieldOptions, m *pilosa.FieldOptions)
m.Min = options.Min m.Min = options.Min
m.Max = options.Max m.Max = options.Max
m.Base = options.Base m.Base = options.Base
m.Scale = options.Scale
m.BitDepth = uint(options.BitDepth) m.BitDepth = uint(options.BitDepth)
m.TimeQuantum = pilosa.TimeQuantum(options.TimeQuantum) m.TimeQuantum = pilosa.TimeQuantum(options.TimeQuantum)
m.Keys = options.Keys m.Keys = options.Keys
m.ForeignIndex = options.ForeignIndex
} }
func decodeNodes(a []*internal.Node, m []*pilosa.Node) { func decodeNodes(a []*internal.Node, m []*pilosa.Node) {
@ -963,6 +1003,10 @@ func decodeQueryRequest(pb *internal.QueryRequest, m *pilosa.QueryRequest) {
m.Remote = pb.Remote m.Remote = pb.Remote
m.ExcludeRowAttrs = pb.ExcludeRowAttrs m.ExcludeRowAttrs = pb.ExcludeRowAttrs
m.ExcludeColumns = pb.ExcludeColumns m.ExcludeColumns = pb.ExcludeColumns
m.EmbeddedData = make([]*pilosa.Row, len(pb.EmbeddedData))
for i := range pb.EmbeddedData {
m.EmbeddedData[i] = decodeRow(pb.EmbeddedData[i])
}
} }
func decodeImportRequest(pb *internal.ImportRequest, m *pilosa.ImportRequest) { func decodeImportRequest(pb *internal.ImportRequest, m *pilosa.ImportRequest) {
@ -983,6 +1027,8 @@ func decodeImportValueRequest(pb *internal.ImportValueRequest, m *pilosa.ImportV
m.ColumnIDs = pb.ColumnIDs m.ColumnIDs = pb.ColumnIDs
m.ColumnKeys = pb.ColumnKeys m.ColumnKeys = pb.ColumnKeys
m.Values = pb.Values m.Values = pb.Values
m.FloatValues = pb.FloatValues
m.StringValues = pb.StringValues
} }
func decodeImportRoaringRequest(pb *internal.ImportRoaringRequest, m *pilosa.ImportRoaringRequest) { func decodeImportRoaringRequest(pb *internal.ImportRoaringRequest, m *pilosa.ImportRoaringRequest) {
@ -994,6 +1040,14 @@ func decodeImportRoaringRequest(pb *internal.ImportRoaringRequest, m *pilosa.Imp
m.Views = views m.Views = views
} }
func decodeImportColumnAttrsRequest(pb *internal.ImportColumnAttrsRequest, m *pilosa.ImportColumnAttrsRequest) {
m.Index = pb.Index
m.Shard = pb.Shard
m.AttrKey = pb.AttrKey
m.AttrVals = pb.AttrVals
m.ColumnIDs = pb.ColumnIDs
}
func decodeImportResponse(pb *internal.ImportResponse, m *pilosa.ImportResponse) { func decodeImportResponse(pb *internal.ImportResponse, m *pilosa.ImportResponse) {
m.Err = pb.Err m.Err = pb.Err
} }
@ -1057,6 +1111,7 @@ const (
queryResultTypeNil uint32 = iota queryResultTypeNil uint32 = iota
queryResultTypeRow queryResultTypeRow
queryResultTypePairs queryResultTypePairs
queryResultTypePairsField
queryResultTypeValCount queryResultTypeValCount
queryResultTypeUint64 queryResultTypeUint64
queryResultTypeBool queryResultTypeBool
@ -1064,14 +1119,20 @@ const (
queryResultTypeGroupCounts queryResultTypeGroupCounts
queryResultTypeRowIdentifiers queryResultTypeRowIdentifiers
queryResultTypePair queryResultTypePair
queryResultTypePairField
queryResultTypeSignedRow
) )
func decodeQueryResult(pb *internal.QueryResult) interface{} { func decodeQueryResult(pb *internal.QueryResult) interface{} {
switch pb.Type { switch pb.Type {
case queryResultTypeSignedRow:
return decodeSignedRow(pb.SignedRow)
case queryResultTypeRow: case queryResultTypeRow:
return decodeRow(pb.Row) return decodeRow(pb.Row)
case queryResultTypePairs: case queryResultTypePairs:
return decodePairs(pb.Pairs) return decodePairs(pb.Pairs)
case queryResultTypePairsField:
return decodePairsField(pb.PairsField)
case queryResultTypeValCount: case queryResultTypeValCount:
return decodeValCount(pb.ValCount) return decodeValCount(pb.ValCount)
case queryResultTypeUint64: case queryResultTypeUint64:
@ -1088,6 +1149,8 @@ func decodeQueryResult(pb *internal.QueryResult) interface{} {
return decodeGroupCounts(pb.GroupCounts) return decodeGroupCounts(pb.GroupCounts)
case queryResultTypePair: case queryResultTypePair:
return decodePair(pb.Pairs[0]) return decodePair(pb.Pairs[0])
case queryResultTypePairField:
return decodePairField(pb.Pairs[0])
} }
panic(fmt.Sprintf("unknown type: %d", pb.Type)) panic(fmt.Sprintf("unknown type: %d", pb.Type))
} }
@ -1095,15 +1158,32 @@ func decodeQueryResult(pb *internal.QueryResult) interface{} {
// DecodeRow converts r from its internal representation. // DecodeRow converts r from its internal representation.
func decodeRow(pr *internal.Row) *pilosa.Row { func decodeRow(pr *internal.Row) *pilosa.Row {
if pr == nil { if pr == nil {
return nil return pilosa.NewRow()
} }
r := pilosa.NewRow() var r *pilosa.Row
r.Attrs = decodeAttrs(pr.Attrs) if len(pr.Roaring) > 0 {
r.Keys = pr.Keys r = pilosa.NewRowFromRoaring(pr.Roaring)
} else {
r = pilosa.NewRow()
for _, v := range pr.Columns { for _, v := range pr.Columns {
r.SetBit(v) r.SetBit(v)
} }
}
r.Attrs = decodeAttrs(pr.Attrs)
r.Keys = pr.Keys
return r
}
func decodeSignedRow(pr *internal.SignedRow) pilosa.SignedRow {
if pr == nil {
return pilosa.SignedRow{}
}
r := pilosa.SignedRow{
Pos: decodeRow(pr.Pos),
Neg: decodeRow(pr.Neg),
}
return r return r
} }
@ -1151,6 +1231,7 @@ func decodeGroupCounts(a []*internal.GroupCount) []pilosa.GroupCount {
other[i] = pilosa.GroupCount{ other[i] = pilosa.GroupCount{
Group: decodeFieldRows(a[i].Group), Group: decodeFieldRows(a[i].Group),
Count: a[i].Count, Count: a[i].Count,
Sum: a[i].Sum,
} }
} }
return other return other
@ -1178,6 +1259,17 @@ func decodePairs(a []*internal.Pair) []pilosa.Pair {
return other return other
} }
func decodePairsField(a *internal.PairsField) *pilosa.PairsField {
other := &pilosa.PairsField{
Pairs: make([]pilosa.Pair, len(a.Pairs)),
}
for i := range a.Pairs {
other.Pairs[i] = decodePair(a.Pairs[i])
}
other.Field = a.Field
return other
}
func decodePair(pb *internal.Pair) pilosa.Pair { func decodePair(pb *internal.Pair) pilosa.Pair {
return pilosa.Pair{ return pilosa.Pair{
ID: pb.ID, ID: pb.ID,
@ -1186,6 +1278,17 @@ func decodePair(pb *internal.Pair) pilosa.Pair {
} }
} }
func decodePairField(pb *internal.Pair) pilosa.PairField {
return pilosa.PairField{
Pair: pilosa.Pair{
ID: pb.ID,
Key: pb.Key,
Count: pb.Count,
},
//Field: pb.Field, // TODO: in order to have this, we need PairField in QueryResponse.
}
}
func decodeValCount(pb *internal.ValCount) pilosa.ValCount { func decodeValCount(pb *internal.ValCount) pilosa.ValCount {
return pilosa.ValCount{ return pilosa.ValCount{
Val: pb.Val, Val: pb.Val,
@ -1209,16 +1312,29 @@ func encodeColumnAttrSet(set *pilosa.ColumnAttrSet) *internal.ColumnAttrSet {
} }
} }
func encodeSignedRow(r pilosa.SignedRow) *internal.SignedRow {
ir := &internal.SignedRow{
Pos: encodeRow(r.Pos),
Neg: encodeRow(r.Neg),
}
return ir
}
func encodeRow(r *pilosa.Row) *internal.Row { func encodeRow(r *pilosa.Row) *internal.Row {
if r == nil { if r == nil {
return nil return nil
} }
return &internal.Row{ ir := &internal.Row{
Columns: r.Columns(),
Keys: r.Keys, Keys: r.Keys,
Attrs: encodeAttrs(r.Attrs), Attrs: encodeAttrs(r.Attrs),
} }
if true {
ir.Columns = r.Columns()
} else {
ir.Roaring = r.Roaring()
}
return ir
} }
func encodeRowIdentifiers(r pilosa.RowIdentifiers) *internal.RowIdentifiers { func encodeRowIdentifiers(r pilosa.RowIdentifiers) *internal.RowIdentifiers {
@ -1235,6 +1351,7 @@ func encodeGroupCounts(counts []pilosa.GroupCount) []*internal.GroupCount {
result[i] = &internal.GroupCount{ result[i] = &internal.GroupCount{
Group: encodeFieldRows(counts[i].Group), Group: encodeFieldRows(counts[i].Group),
Count: counts[i].Count, Count: counts[i].Count,
Sum: counts[i].Sum,
} }
} }
return result return result
@ -1267,6 +1384,17 @@ func encodePairs(a pilosa.Pairs) []*internal.Pair {
return other return other
} }
func encodePairsField(a *pilosa.PairsField) *internal.PairsField {
other := &internal.PairsField{
Pairs: make([]*internal.Pair, len(a.Pairs)),
}
for i := range a.Pairs {
other.Pairs[i] = encodePair(a.Pairs[i])
}
other.Field = a.Field
return other
}
func encodePair(p pilosa.Pair) *internal.Pair { func encodePair(p pilosa.Pair) *internal.Pair {
return &internal.Pair{ return &internal.Pair{
ID: p.ID, ID: p.ID,
@ -1275,6 +1403,17 @@ func encodePair(p pilosa.Pair) *internal.Pair {
} }
} }
func encodePairField(p pilosa.PairField) *internal.Pair {
/*
// TODO: in order to have this, we need PairField in QueryResponse.
return &internal.Pair{
Pair: encodePair(p.Pair),
Field: p.Field,
}
*/
return encodePair(p.Pair)
}
func encodeValCount(vc pilosa.ValCount) *internal.ValCount { func encodeValCount(vc pilosa.ValCount) *internal.ValCount {
return &internal.ValCount{ return &internal.ValCount{
Val: vc.Val, Val: vc.Val,

File diff suppressed because it is too large Load diff

View file

@ -46,7 +46,7 @@ func TestExecutor_TranslateGroupByCall(t *testing.T) {
t.Fatalf("creating fields %v, %v, %v", erra, errb, errc) t.Fatalf("creating fields %v, %v, %v", erra, errb, errc)
} }
query, err := pql.ParseString(`GroupBy(Rows(ak), Rows(b), Rows(ck), previous=["la", 0, "ha"])`) query, err := pql.ParseString(`GroupBy(Rows(ak), Rows(b), Rows(ck), previous=["la", 0, "ha"], having=Condition(count > 10))`)
if err != nil { if err != nil {
t.Fatalf("parsing query: %v", err) t.Fatalf("parsing query: %v", err)
} }
@ -64,6 +64,19 @@ func TestExecutor_TranslateGroupByCall(t *testing.T) {
} }
} }
if having, hok := c.Args["having"].(*pql.Call); !hok {
t.Fatal("expected having to be a call")
} else if cond, cok := having.Args["count"].(*pql.Condition); !cok {
t.Fatal("expected condition to be a count")
} else if cond.Op != pql.GT {
t.Fatal("expected condition op to be >")
} else {
val, ok := cond.Uint64Value()
if !ok || val != uint64(10) {
t.Fatal("expected condition val to be uint64(10)")
}
}
errTests := []struct { errTests := []struct {
pql string pql string
err string err string
@ -222,3 +235,144 @@ func TestFieldRowMarshalJSON(t *testing.T) {
t.Fatalf("unexpected json: %s", b) t.Fatalf("unexpected json: %s", b)
} }
} }
func TestExecutor_GroupCountCondition(t *testing.T) {
t.Run("satisfiesCondition", func(t *testing.T) {
type condCheck struct {
cond string
exp bool
}
tests := []struct {
groupCount GroupCount
checks []condCheck
}{
{
groupCount: GroupCount{Count: 100},
checks: []condCheck{
{cond: "count == 99", exp: false},
{cond: "count != 99", exp: true},
{cond: "count < 99", exp: false},
{cond: "count <= 99", exp: false},
{cond: "count > 99", exp: true},
{cond: "count >= 99", exp: true},
{cond: "count == 100", exp: true},
{cond: "count != 100", exp: false},
{cond: "count < 100", exp: false},
{cond: "count <= 100", exp: true},
{cond: "count > 100", exp: false},
{cond: "count >= 100", exp: true},
{cond: "count == 101", exp: false},
{cond: "count != 101", exp: true},
{cond: "count < 101", exp: true},
{cond: "count <= 101", exp: true},
{cond: "count > 101", exp: false},
{cond: "count >= 101", exp: false},
{cond: "98 < count < 100", exp: false},
{cond: "98 < count <= 100", exp: true},
{cond: "98 < count < 101", exp: true},
{cond: "100 <= count < 102", exp: true},
{cond: "100 < count < 102", exp: false},
{cond: "98 <= count <= 102", exp: true},
},
},
{
groupCount: GroupCount{Sum: 100},
checks: []condCheck{
{cond: "sum == 99", exp: false},
{cond: "sum != 99", exp: true},
{cond: "sum < 99", exp: false},
{cond: "sum <= 99", exp: false},
{cond: "sum > 99", exp: true},
{cond: "sum >= 99", exp: true},
{cond: "sum == 100", exp: true},
{cond: "sum != 100", exp: false},
{cond: "sum < 100", exp: false},
{cond: "sum <= 100", exp: true},
{cond: "sum > 100", exp: false},
{cond: "sum >= 100", exp: true},
{cond: "sum == 101", exp: false},
{cond: "sum != 101", exp: true},
{cond: "sum < 101", exp: true},
{cond: "sum <= 101", exp: true},
{cond: "sum > 101", exp: false},
{cond: "sum >= 101", exp: false},
{cond: "98 < sum < 100", exp: false},
{cond: "98 < sum <= 100", exp: true},
{cond: "98 < sum < 101", exp: true},
{cond: "100 <= sum < 102", exp: true},
{cond: "100 < sum < 102", exp: false},
{cond: "98 <= sum <= 102", exp: true},
},
},
{
groupCount: GroupCount{Sum: -100},
checks: []condCheck{
{cond: "sum == -99", exp: false},
{cond: "sum != -99", exp: true},
{cond: "sum < -99", exp: true},
{cond: "sum <= -99", exp: true},
{cond: "sum > -99", exp: false},
{cond: "sum >= -99", exp: false},
{cond: "sum == -100", exp: true},
{cond: "sum != -100", exp: false},
{cond: "sum < -100", exp: false},
{cond: "sum <= -100", exp: true},
{cond: "sum > -100", exp: false},
{cond: "sum >= -100", exp: true},
{cond: "sum == -101", exp: false},
{cond: "sum != -101", exp: true},
{cond: "sum < -101", exp: false},
{cond: "sum <= -101", exp: false},
{cond: "sum > -101", exp: true},
{cond: "sum >= -101", exp: true},
{cond: "-100 < sum < -98", exp: false},
{cond: "-100 <= sum < -98", exp: true},
{cond: "-101 < sum < -98", exp: true},
{cond: "-102 < sum <= -100", exp: true},
{cond: "-102 < sum < -100", exp: false},
{cond: "-102 <= sum <= -98", exp: true},
},
},
}
for i, test := range tests {
t.Run(fmt.Sprintf("test (#%d):", i), func(t *testing.T) {
for j, check := range test.checks {
t.Run(fmt.Sprintf("check (#%d):", j), func(t *testing.T) {
query, err := pql.ParseString(fmt.Sprintf("GroupBy(Rows(a), having=Condition(%s))", check.cond))
if err != nil {
t.Fatalf("parsing query: %v", err)
}
c := query.Calls[0]
having := c.Args["having"].(*pql.Call)
var got bool
for subj, cond := range having.Args {
switch subj {
case "count", "sum":
condition, ok := cond.(*pql.Condition)
if !ok {
t.Fatalf("not a valid condition")
}
got = test.groupCount.satisfiesCondition(subj, condition)
}
}
if got != check.exp {
t.Fatalf("expected: %v, but got: %v", check.exp, got)
}
})
}
})
}
})
}

View file

@ -526,22 +526,18 @@ func TestExecutor_Execute_Set(t *testing.T) {
if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(1, f=11)`}); err != nil { if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(1, f=11)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else { } else if !res.Results[0].(bool) {
if !res.Results[0].(bool) {
t.Fatalf("expected column changed") t.Fatalf("expected column changed")
} }
}
if n := hldr.Row("i", "f", 11).Count(); n != 1 { if n := hldr.Row("i", "f", 11).Count(); n != 1 {
t.Fatalf("unexpected row count: %d", n) t.Fatalf("unexpected row count: %d", n)
} }
if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(1, f=11)`}); err != nil { if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(1, f=11)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else { } else if res.Results[0].(bool) {
if res.Results[0].(bool) {
t.Fatalf("expected column unchanged") t.Fatalf("expected column unchanged")
} }
}
}) })
t.Run("ErrInvalidColValueType", func(t *testing.T) { t.Run("ErrInvalidColValueType", func(t *testing.T) {
@ -591,21 +587,29 @@ func TestExecutor_Execute_Set(t *testing.T) {
if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set("foo", f=11)`}); err != nil { if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set("foo", f=11)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else { } else if !res.Results[0].(bool) {
if !res.Results[0].(bool) {
t.Fatalf("expected column changed") t.Fatalf("expected column changed")
} }
}
if n := hldr.Row("i", "f", 11).Count(); n != 1 { if n := hldr.Row("i", "f", 11).Count(); n != 1 {
t.Fatalf("unexpected row count: %d", n) t.Fatalf("unexpected row count: %d", n)
} }
if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set("foo", f=11)`}); err != nil { if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set("foo", f=11)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else { } else if res.Results[0].(bool) {
if res.Results[0].(bool) {
t.Fatalf("expected column unchanged") t.Fatalf("expected column unchanged")
} }
if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(2, f=11)`}); err != nil {
t.Fatal(err)
} else if !res.Results[0].(bool) {
t.Fatalf("expected column changed with integer column key")
}
if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(2, f=11)`}); err != nil {
t.Fatal(err)
} else if res.Results[0].(bool) {
t.Fatalf("expected column unchanged with integer column key")
} }
}) })
@ -617,9 +621,15 @@ func TestExecutor_Execute_Set(t *testing.T) {
t.Fatal(err) t.Fatal(err)
} }
if _, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(2, f=1)`}); err == nil || errors.Cause(err).Error() != `column value must be a string when index 'keys' option enabled` { if _, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(2.1, f=1)`}); err == nil || strings.Contains(err.Error(), `column value must be a string or non-negative integer`) {
t.Fatal(err) t.Fatal(err)
} }
if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(2, f=1)`}); err != nil {
t.Fatal(err)
} else if !res.Results[0].(bool) {
t.Fatalf("expected column changed with integer column key")
}
}) })
t.Run("ErrInvalidRowValueType", func(t *testing.T) { t.Run("ErrInvalidRowValueType", func(t *testing.T) {
@ -627,9 +637,16 @@ func TestExecutor_Execute_Set(t *testing.T) {
if _, err := index.CreateField("f", pilosa.OptFieldTypeDefault(), pilosa.OptFieldKeys()); err != nil { if _, err := index.CreateField("f", pilosa.OptFieldTypeDefault(), pilosa.OptFieldKeys()); err != nil {
t.Fatal(err) t.Fatal(err)
} }
if _, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "inokey", Query: `Set(2, f=1)`}); err == nil || errors.Cause(err).Error() != `row value must be a string when field 'keys' option enabled` { if _, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "inokey", Query: `Set(2, f=1.2)`}); err == nil || !strings.Contains(err.Error(), "row value must be a string or non-negative integer") {
t.Fatal(err) t.Fatal(err)
} }
if res, err := cmd.API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(2, f=9)`}); err != nil {
t.Fatal(err)
} else if !res.Results[0].(bool) {
t.Fatalf("expected column changed with integer column key")
}
}) })
}) })
} }
@ -928,9 +945,12 @@ func TestExecutor_Execute_TopN(t *testing.T) {
if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=2)`}); err != nil { if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=2)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(result.Results[0], []pilosa.Pair{ } else if !reflect.DeepEqual(result.Results[0], &pilosa.PairsField{
Pairs: []pilosa.Pair{
{ID: 0, Count: 5}, {ID: 0, Count: 5},
{ID: 10, Count: 2}, {ID: 10, Count: 2},
},
Field: "f",
}) { }) {
t.Fatalf("unexpected result: %s", spew.Sdump(result)) t.Fatalf("unexpected result: %s", spew.Sdump(result))
} }
@ -969,9 +989,12 @@ func TestExecutor_Execute_TopN(t *testing.T) {
if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=2)`}); err != nil { if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=2)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(result.Results[0], []pilosa.Pair{ } else if !reflect.DeepEqual(result.Results[0], &pilosa.PairsField{
Pairs: []pilosa.Pair{
{ID: 0, Count: 5}, {ID: 0, Count: 5},
{ID: 10, Count: 2}, {ID: 10, Count: 2},
},
Field: "f",
}) { }) {
t.Fatalf("unexpected result: %s", spew.Sdump(result)) t.Fatalf("unexpected result: %s", spew.Sdump(result))
} }
@ -1010,12 +1033,17 @@ func TestExecutor_Execute_TopN(t *testing.T) {
if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=2)`}); err != nil { if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=2)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(result.Results[0], []pilosa.Pair{ } else {
if !reflect.DeepEqual(result.Results[0], &pilosa.PairsField{
Pairs: []pilosa.Pair{
{Key: "zero", Count: 5}, {Key: "zero", Count: 5},
{Key: "ten", Count: 2}, {Key: "ten", Count: 2},
},
Field: "f",
}) { }) {
t.Fatalf("unexpected result: %s", spew.Sdump(result)) t.Fatalf("unexpected result: %s", spew.Sdump(result))
} }
}
}) })
t.Run("RowKeyColumnKey", func(t *testing.T) { t.Run("RowKeyColumnKey", func(t *testing.T) {
@ -1052,10 +1080,13 @@ func TestExecutor_Execute_TopN(t *testing.T) {
if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=2)`}); err != nil { if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=2)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if diff := cmp.Diff(result.Results, []interface{}{ } else if diff := cmp.Diff(result.Results, []interface{}{
[]pilosa.Pair{ &pilosa.PairsField{
Pairs: []pilosa.Pair{
{Key: "foo", Count: 5}, {Key: "foo", Count: 5},
{Key: "bar", Count: 2}, {Key: "bar", Count: 2},
}, },
Field: "f",
},
}); diff != "" { }); diff != "" {
t.Fatal(diff) t.Fatal(diff)
} }
@ -1137,8 +1168,11 @@ func TestExecutor_Execute_TopN_fill(t *testing.T) {
// Execute query. // Execute query.
if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=1)`}); err != nil { if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=1)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(result.Results, []interface{}{[]pilosa.Pair{ } else if !reflect.DeepEqual(result.Results, []interface{}{&pilosa.PairsField{
Pairs: []pilosa.Pair{
{ID: 0, Count: 4}, {ID: 0, Count: 4},
},
Field: "f",
}}) { }}) {
t.Fatalf("unexpected result: %s", spew.Sdump(result)) t.Fatalf("unexpected result: %s", spew.Sdump(result))
} }
@ -1171,8 +1205,11 @@ func TestExecutor_Execute_TopN_fill_small(t *testing.T) {
// Execute query. // Execute query.
if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=1)`}); err != nil { if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=1)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(result.Results, []interface{}{[]pilosa.Pair{ } else if !reflect.DeepEqual(result.Results, []interface{}{&pilosa.PairsField{
Pairs: []pilosa.Pair{
{ID: 0, Count: 5}, {ID: 0, Count: 5},
},
Field: "f",
}}) { }}) {
t.Fatalf("unexpected result: %s", spew.Sdump(result)) t.Fatalf("unexpected result: %s", spew.Sdump(result))
} }
@ -1207,10 +1244,13 @@ func TestExecutor_Execute_TopN_Src(t *testing.T) {
// Execute query. // Execute query.
if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, Row(other=100), n=3)`}); err != nil { if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, Row(other=100), n=3)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(result.Results, []interface{}{[]pilosa.Pair{ } else if !reflect.DeepEqual(result.Results, []interface{}{&pilosa.PairsField{
Pairs: []pilosa.Pair{
{ID: 20, Count: 3}, {ID: 20, Count: 3},
{ID: 10, Count: 2}, {ID: 10, Count: 2},
{ID: 0, Count: 1}, {ID: 0, Count: 1},
},
Field: "f",
}}) { }}) {
t.Fatalf("unexpected result: %s", spew.Sdump(result)) t.Fatalf("unexpected result: %s", spew.Sdump(result))
} }
@ -1230,8 +1270,11 @@ func TestExecutor_Execute_TopN_Attr(t *testing.T) {
} }
if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=1, attrName="category", attrValues=[123])`}); err != nil { if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=1, attrName="category", attrValues=[123])`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(result.Results, []interface{}{[]pilosa.Pair{ } else if !reflect.DeepEqual(result.Results, []interface{}{&pilosa.PairsField{
Pairs: []pilosa.Pair{
{ID: 10, Count: 1}, {ID: 10, Count: 1},
},
Field: "f",
}}) { }}) {
t.Fatalf("unexpected result: %s", spew.Sdump(result)) t.Fatalf("unexpected result: %s", spew.Sdump(result))
} }
@ -1253,8 +1296,11 @@ func TestExecutor_Execute_TopN_Attr_Src(t *testing.T) {
} }
if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, Row(f=10), n=1, attrName="category", attrValues=[123])`}); err != nil { if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, Row(f=10), n=1, attrName="category", attrValues=[123])`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(result.Results, []interface{}{[]pilosa.Pair{ } else if !reflect.DeepEqual(result.Results, []interface{}{&pilosa.PairsField{
Pairs: []pilosa.Pair{
{ID: 10, Count: 1}, {ID: 10, Count: 1},
},
Field: "f",
}}) { }}) {
t.Fatalf("unexpected result: %s", spew.Sdump(result)) t.Fatalf("unexpected result: %s", spew.Sdump(result))
} }
@ -1448,7 +1494,10 @@ func TestExecutor_Execute_MinMaxRow(t *testing.T) {
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
target := pilosa.Pair{ID: 1, Count: 1} target := pilosa.PairField{
Pair: pilosa.Pair{ID: 1, Count: 1},
Field: "f",
}
if !reflect.DeepEqual(target, result.Results[0]) { if !reflect.DeepEqual(target, result.Results[0]) {
t.Fatalf("unexpected result %v != %v", target, result.Results[0]) t.Fatalf("unexpected result %v != %v", target, result.Results[0])
} }
@ -1459,7 +1508,10 @@ func TestExecutor_Execute_MinMaxRow(t *testing.T) {
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
target := pilosa.Pair{ID: 10000, Count: 1} target := pilosa.PairField{
Pair: pilosa.Pair{ID: 10000, Count: 1},
Field: "f",
}
if !reflect.DeepEqual(target, result.Results[0]) { if !reflect.DeepEqual(target, result.Results[0]) {
t.Fatalf("unexpected result %v != %v", target, result.Results[0]) t.Fatalf("unexpected result %v != %v", target, result.Results[0])
} }
@ -1495,7 +1547,10 @@ func TestExecutor_Execute_MinMaxRow(t *testing.T) {
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
target := pilosa.Pair{Key: "seven-thousand", ID: 1, Count: 1} target := pilosa.PairField{
Pair: pilosa.Pair{Key: "seven-thousand", ID: 1, Count: 1},
Field: "f",
}
if !reflect.DeepEqual(target, result.Results[0]) { if !reflect.DeepEqual(target, result.Results[0]) {
t.Fatalf("unexpected result %v != %v", target, result.Results[0]) t.Fatalf("unexpected result %v != %v", target, result.Results[0])
} }
@ -1506,7 +1561,10 @@ func TestExecutor_Execute_MinMaxRow(t *testing.T) {
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
target := pilosa.Pair{Key: "five-thousand", ID: 5, Count: 1} target := pilosa.PairField{
Pair: pilosa.Pair{Key: "five-thousand", ID: 5, Count: 1},
Field: "f",
}
if !reflect.DeepEqual(target, result.Results[0]) { if !reflect.DeepEqual(target, result.Results[0]) {
t.Fatalf("unexpected result %v != %v", target, result.Results[0]) t.Fatalf("unexpected result %v != %v", target, result.Results[0])
} }
@ -2403,7 +2461,6 @@ func TestExecutor_Execute_Remote_Row(t *testing.T) {
if err != nil { if err != nil {
t.Fatalf("creating field: %v", err) t.Fatalf("creating field: %v", err)
} }
if _, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: ` if _, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `
Set(500001, fn=5) Set(500001, fn=5)
Set(1500001, fn=5) Set(1500001, fn=5)
@ -2415,11 +2472,11 @@ Set(3500003, fn=3)
Set(500001, fn=4) Set(500001, fn=4)
Set(4500001, fn=4) Set(4500001, fn=4)
`}); err != nil { `}); err != nil {
t.Fatalf("quuerying remote: %v", err) t.Fatalf("querying remote: %v", err)
} }
err := c[0].API.RecalculateCaches(context.Background()) err := c[0].API.RecalculateCaches(context.Background())
if err != nil { if err != nil {
t.Fatalf("recalcing caches: %v", err) t.Fatalf("recalculating caches: %v", err)
} }
if res, err := c[1].API.Query(context.Background(), &pilosa.QueryRequest{ if res, err := c[1].API.Query(context.Background(), &pilosa.QueryRequest{
@ -2427,10 +2484,13 @@ Set(4500001, fn=4)
Query: `TopN(fn, n=3)`, Query: `TopN(fn, n=3)`,
}); err != nil { }); err != nil {
t.Fatalf("topn querying: %v", err) t.Fatalf("topn querying: %v", err)
} else if !reflect.DeepEqual(res.Results, []interface{}{[]pilosa.Pair{ } else if !reflect.DeepEqual(res.Results, []interface{}{&pilosa.PairsField{
Pairs: []pilosa.Pair{
{ID: 5, Count: 4}, {ID: 5, Count: 4},
{ID: 3, Count: 3}, {ID: 3, Count: 3},
{ID: 4, Count: 2}, {ID: 4, Count: 2},
},
Field: "fn",
}}) { }}) {
t.Fatalf("topn wrong results: %v", res.Results) t.Fatalf("topn wrong results: %v", res.Results)
} }
@ -2838,6 +2898,175 @@ func TestExecutor_Execute_Not(t *testing.T) {
}) })
} }
// Ensure an all query can be executed.
func TestExecutor_Execute_All(t *testing.T) {
t.Run("ColumnID", func(t *testing.T) {
c := test.MustRunCluster(t, 1)
defer c.Close()
hldr := test.Holder{Holder: c[0].Server.Holder()}
index := hldr.MustCreateIndexIfNotExists("i", pilosa.IndexOptions{TrackExistence: true})
fld, err := index.CreateField("f", pilosa.OptFieldTypeDefault())
if err != nil {
t.Fatal(err)
}
// Create an import request that sets a full shard,
// plus a couple bits set on either side of it, and
// a final bit set in a fourth shard.
//
// shard0 shard1 shard2 shard3
// |----------|----------|----------|----------|
// | **|**********|** | *
//
bitCount := ShardWidth + 5
req := &pilosa.ImportRequest{
Index: index.Name(),
Field: fld.Name(),
Shard: 0,
RowIDs: make([]uint64, bitCount),
ColumnIDs: make([]uint64, bitCount),
}
for i := 0; i < bitCount-1; i++ {
req.RowIDs[i] = 10
req.ColumnIDs[i] = uint64(i + ShardWidth - 2)
}
req.RowIDs[bitCount-1] = 10
req.ColumnIDs[bitCount-1] = uint64((3 * ShardWidth) + 2)
if err := c[0].API.Import(context.Background(), req); err != nil {
t.Fatal(err)
}
tests := []struct {
qry string
expCols []uint64
expCnt uint64
}{
{qry: "All()", expCols: req.ColumnIDs, expCnt: uint64(bitCount)},
{qry: "All(limit=1)", expCols: req.ColumnIDs[:1], expCnt: 1},
{qry: "All(limit=4)", expCols: req.ColumnIDs[:4], expCnt: 4},
{qry: "All(limit=4, offset=4)", expCols: req.ColumnIDs[4:8], expCnt: 4},
{qry: fmt.Sprintf("All(limit=4, offset=%d)", bitCount-5), expCols: req.ColumnIDs[bitCount-5 : bitCount-1], expCnt: 4},
{qry: fmt.Sprintf("All(limit=1, offset=%d)", bitCount-2), expCols: req.ColumnIDs[bitCount-2 : bitCount-1], expCnt: 1},
{qry: fmt.Sprintf("All(limit=1, offset=%d)", bitCount-2), expCols: req.ColumnIDs[bitCount-2 : bitCount-1], expCnt: 1},
{qry: fmt.Sprintf("All(limit=4, offset=%d)", bitCount-2), expCols: req.ColumnIDs[bitCount-2:], expCnt: 2},
{qry: fmt.Sprintf("All(limit=4, offset=%d)", bitCount+1), expCols: []uint64{}, expCnt: 0},
{qry: fmt.Sprintf("All(limit=2, offset=%d)", bitCount-3), expCols: req.ColumnIDs[bitCount-3 : bitCount-1], expCnt: 2},
{qry: fmt.Sprintf("All(limit=2, offset=%d)", bitCount-5), expCols: req.ColumnIDs[bitCount-5 : bitCount-3], expCnt: 2},
{qry: "All(limit=2, offset=2)", expCols: req.ColumnIDs[2:4], expCnt: 2},
{qry: "All(limit=1, offset=1)", expCols: req.ColumnIDs[1:2], expCnt: 1},
{qry: fmt.Sprintf("All(limit=%d, offset=2)", ShardWidth), expCols: req.ColumnIDs[2 : bitCount-3], expCnt: ShardWidth},
}
for i, test := range tests {
if res, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: test.qry}); err != nil {
t.Fatal(err)
} else if cnt := res.Results[0].(*pilosa.Row).Count(); cnt != test.expCnt {
t.Fatalf("test %d, unexpected count, got: %d, but expected: %d", i, cnt, test.expCnt)
} else if cols := res.Results[0].(*pilosa.Row).Columns(); !reflect.DeepEqual(cols, test.expCols) {
// If the error results are too large, just show the count.
if len(cols) > 1000 || len(test.expCols) > 1000 {
t.Fatalf("test %d, unexpected columns, got: len(%d), but expected: len(%d)", i, len(cols), len(test.expCols))
} else {
t.Fatalf("test %d, unexpected columns, got: %v, but expected: %v", i, cols, test.expCols)
}
}
}
})
t.Run("ColumnKey", func(t *testing.T) {
c := test.MustRunCluster(t, 1, []server.CommandOption{
server.OptCommandServerOptions(
pilosa.OptServerOpenTranslateStore(boltdb.OpenTranslateStore),
pilosa.OptServerOpenTranslateReader(http.GetOpenTranslateReaderFunc(nil)),
),
})
defer c.Close()
hldr := test.Holder{Holder: c[0].Server.Holder()}
index := hldr.MustCreateIndexIfNotExists("i", pilosa.IndexOptions{TrackExistence: true, Keys: true})
fld, err := index.CreateField("f", pilosa.OptFieldTypeDefault())
if err != nil {
t.Fatal(err)
}
// Create an import request that sets key columns
//
// shard0
// |----------|
// |**** |
//
bitCount := 4
req := &pilosa.ImportRequest{
Index: index.Name(),
Field: fld.Name(),
Shard: 0,
RowIDs: make([]uint64, bitCount),
ColumnKeys: make([]string, bitCount),
}
for i := 0; i < bitCount; i++ {
req.RowIDs[i] = 10
req.ColumnKeys[i] = fmt.Sprintf("c%d", i)
}
if err := c[0].API.Import(context.Background(), req); err != nil {
t.Fatal(err)
}
tests := []struct {
qry string
expCols []string
expCnt uint64
}{
{qry: "All()", expCols: req.ColumnKeys, expCnt: uint64(bitCount)},
{qry: "All(limit=1)", expCols: req.ColumnKeys[:1], expCnt: 1},
{qry: "All(limit=4)", expCols: req.ColumnKeys, expCnt: 4},
{qry: "All(limit=5)", expCols: req.ColumnKeys, expCnt: 4},
{qry: "All(limit=1, offset=1)", expCols: req.ColumnKeys[1:2], expCnt: 1},
{qry: "All(limit=4, offset=1)", expCols: req.ColumnKeys[1:], expCnt: 3},
{qry: "All(limit=4, offset=5)", expCols: nil, expCnt: 0},
}
for i, test := range tests {
if res, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: test.qry}); err != nil {
t.Fatal(err)
} else if cnt := len(res.Results[0].(*pilosa.Row).Keys); uint64(cnt) != test.expCnt {
t.Fatalf("test %d, unexpected count, got: %d, but expected: %d", i, cnt, test.expCnt)
} else if cols := res.Results[0].(*pilosa.Row).Keys; !reflect.DeepEqual(cols, test.expCols) {
// If the error results are too large, just show the count.
if len(cols) > 1000 || len(test.expCols) > 1000 {
t.Fatalf("test %d, unexpected columns, got: len(%d), but expected: len(%d)", i, len(cols), len(test.expCols))
} else {
t.Fatalf("test %d, unexpected columns, got: %T, but expected: %T", i, cols, test.expCols)
}
}
}
})
// Ensure that a query which uses All() at the shard level can call it.
t.Run("AllShard", func(t *testing.T) {
c := test.MustRunCluster(t, 1)
defer c.Close()
hldr := test.Holder{Holder: c[0].Server.Holder()}
index := hldr.MustCreateIndexIfNotExists("i", pilosa.IndexOptions{TrackExistence: true})
_, err := index.CreateField("f", pilosa.OptFieldTypeDefault())
if err != nil {
t.Fatal(err)
}
if _, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `
Set(3001, f=3)
Set(5001, f=5)
Set(5002, f=5)
`}); err != nil {
t.Fatalf("querying remote: %v", err)
}
expCols := []uint64{5001, 5002}
if res, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: "Intersect(All(), Row(f=5))"}); err != nil {
t.Fatal(err)
} else if cols := res.Results[0].(*pilosa.Row).Columns(); !reflect.DeepEqual(cols, expCols) {
t.Fatalf("unexpected columns, got: %v, but expected: %v", cols, expCols)
}
})
}
// Ensure a row can be cleared. // Ensure a row can be cleared.
func TestExecutor_Execute_ClearRow(t *testing.T) { func TestExecutor_Execute_ClearRow(t *testing.T) {
// Set and Mutex tests use the same data and queries // Set and Mutex tests use the same data and queries
@ -3022,10 +3251,13 @@ func TestExecutor_Execute_ClearRow(t *testing.T) {
// Check the TopN results. // Check the TopN results.
if res, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=5)`}); err != nil { if res, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=5)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(res.Results, []interface{}{[]pilosa.Pair{ } else if !reflect.DeepEqual(res.Results, []interface{}{&pilosa.PairsField{
Pairs: []pilosa.Pair{
{ID: 1, Count: 7}, {ID: 1, Count: 7},
{ID: 2, Count: 6}, {ID: 2, Count: 6},
{ID: 3, Count: 5}, {ID: 3, Count: 5},
},
Field: "f",
}}) { }}) {
t.Fatalf("topn wrong results: %v", res.Results) t.Fatalf("topn wrong results: %v", res.Results)
} }
@ -3040,9 +3272,12 @@ func TestExecutor_Execute_ClearRow(t *testing.T) {
// Ensure that the cleared row doesn't show up in TopN (i.e. it was removed from the cache). // Ensure that the cleared row doesn't show up in TopN (i.e. it was removed from the cache).
if res, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=5)`}); err != nil { if res, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `TopN(f, n=5)`}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(res.Results, []interface{}{[]pilosa.Pair{ } else if !reflect.DeepEqual(res.Results, []interface{}{&pilosa.PairsField{
Pairs: []pilosa.Pair{
{ID: 1, Count: 7}, {ID: 1, Count: 7},
{ID: 3, Count: 5}, {ID: 3, Count: 5},
},
Field: "f",
}}) { }}) {
t.Fatalf("topn wrong results: %v", res.Results) t.Fatalf("topn wrong results: %v", res.Results)
} }
@ -3263,30 +3498,40 @@ func TestExecutor_Execute_Rows(t *testing.T) {
}) })
rows := c.Query(t, "i", `Rows(general)`).Results[0].(pilosa.RowIdentifiers) rows := c.Query(t, "i", `Rows(general)`).Results[0].(pilosa.RowIdentifiers)
if !reflect.DeepEqual(rows, pilosa.RowIdentifiers{Rows: []uint64{10, 11, 12, 13}}) { if !reflect.DeepEqual(rows.Rows, []uint64{10, 11, 12, 13}) {
t.Fatalf("unexpected rows: %+v", rows) t.Fatalf("unexpected rows: %+v", rows.Rows)
} else if rows.Keys != nil {
t.Fatalf("unexpected keys: %+v", rows.Keys)
} }
// backwards compatibility // backwards compatibility
// TODO: remove at Pilosa 2.0 // TODO: remove at Pilosa 2.0
rows = c.Query(t, "i", `Rows(field=general)`).Results[0].(pilosa.RowIdentifiers) rows = c.Query(t, "i", `Rows(field=general)`).Results[0].(pilosa.RowIdentifiers)
if !reflect.DeepEqual(rows, pilosa.RowIdentifiers{Rows: []uint64{10, 11, 12, 13}}) { if !reflect.DeepEqual(rows.Rows, []uint64{10, 11, 12, 13}) {
t.Fatalf("unexpected rows: %+v", rows) t.Fatalf("unexpected rows: %+v", rows.Rows)
} else if rows.Keys != nil {
t.Fatalf("unexpected keys: %+v", rows.Keys)
} }
rows = c.Query(t, "i", `Rows(general, limit=2)`).Results[0].(pilosa.RowIdentifiers) rows = c.Query(t, "i", `Rows(general, limit=2)`).Results[0].(pilosa.RowIdentifiers)
if !reflect.DeepEqual(rows, pilosa.RowIdentifiers{Rows: []uint64{10, 11}}) { if !reflect.DeepEqual(rows.Rows, []uint64{10, 11}) {
t.Fatalf("unexpected rows: %+v", rows) t.Fatalf("unexpected rows: %+v", rows.Rows)
} else if rows.Keys != nil {
t.Fatalf("unexpected keys: %+v", rows.Keys)
} }
rows = c.Query(t, "i", `Rows(general, previous=10,limit=2)`).Results[0].(pilosa.RowIdentifiers) rows = c.Query(t, "i", `Rows(general, previous=10,limit=2)`).Results[0].(pilosa.RowIdentifiers)
if !reflect.DeepEqual(rows, pilosa.RowIdentifiers{Rows: []uint64{11, 12}}) { if !reflect.DeepEqual(rows.Rows, []uint64{11, 12}) {
t.Fatalf("unexpected rows: %+v", rows) t.Fatalf("unexpected rows: %+v", rows.Rows)
} else if rows.Keys != nil {
t.Fatalf("unexpected keys: %+v", rows.Keys)
} }
rows = c.Query(t, "i", `Rows(general, column=2)`).Results[0].(pilosa.RowIdentifiers) rows = c.Query(t, "i", `Rows(general, column=2)`).Results[0].(pilosa.RowIdentifiers)
if !reflect.DeepEqual(rows, pilosa.RowIdentifiers{Rows: []uint64{11, 12}}) { if !reflect.DeepEqual(rows.Rows, []uint64{11, 12}) {
t.Fatalf("unexpected rows: %+v", rows) t.Fatalf("unexpected rows: %+v", rows.Rows)
} else if rows.Keys != nil {
t.Fatalf("unexpected keys: %+v", rows.Keys)
} }
} }
@ -3396,15 +3641,25 @@ func TestExecutor_GroupByStrings(t *testing.T) {
c := test.MustRunCluster(t, 1) c := test.MustRunCluster(t, 1)
defer c.Close() defer c.Close()
c.CreateField(t, "istring", pilosa.IndexOptions{Keys: true}, "generals", pilosa.OptFieldKeys()) c.CreateField(t, "istring", pilosa.IndexOptions{Keys: true}, "generals", pilosa.OptFieldKeys())
c.CreateField(t, "istring", pilosa.IndexOptions{Keys: true}, "v", pilosa.OptFieldTypeInt(0, 1000))
req := &pilosa.ImportRequest{ if err := c[0].API.Import(context.Background(), &pilosa.ImportRequest{
Index: "istring", Index: "istring",
Field: "generals", Field: "generals",
Shard: 0, Shard: 0,
RowKeys: []string{"r1", "r2", "r1", "r2", "r1", "r2", "r1", "r2", "r1", "r2"}, RowKeys: []string{"r1", "r2", "r1", "r2", "r1", "r2", "r1", "r2", "r1", "r2"},
ColumnKeys: []string{"c1", "c2", "c3", "c4", "c5", "c6", "c7", "c8", "c9", "c10"}, ColumnKeys: []string{"c1", "c2", "c3", "c4", "c5", "c6", "c7", "c8", "c9", "c10"},
}); err != nil {
t.Fatalf("importing: %v", err)
} }
if err := c[0].API.Import(context.Background(), req); err != nil {
if err := c[0].API.ImportValue(context.Background(), &pilosa.ImportValueRequest{
Index: "istring",
Field: "v",
Shard: 0,
ColumnKeys: []string{"c1", "c2", "c3", "c4", "c5", "c6", "c7", "c8", "c9", "c10"},
Values: []int64{1, 2, 3, 4, 5, 6, 7, 8, 9, 10},
}); err != nil {
t.Fatalf("importing: %v", err) t.Fatalf("importing: %v", err)
} }
@ -3425,6 +3680,29 @@ func TestExecutor_GroupByStrings(t *testing.T) {
{Group: []pilosa.FieldRow{{Field: "generals", RowID: 2, RowKey: "r2"}}, Count: 5}, {Group: []pilosa.FieldRow{{Field: "generals", RowID: 2, RowKey: "r2"}}, Count: 5},
}, },
}, },
{
query: "GroupBy(Rows(generals), aggregate=Sum(field=v))",
expected: []pilosa.GroupCount{
{Group: []pilosa.FieldRow{{Field: "generals", RowID: 1, RowKey: "r1"}}, Count: 5, Sum: 25},
{Group: []pilosa.FieldRow{{Field: "generals", RowID: 2, RowKey: "r2"}}, Count: 5, Sum: 30},
},
},
{
query: "GroupBy(Rows(generals), aggregate=Sum(field=v), having=Condition(sum>25))",
expected: []pilosa.GroupCount{
{Group: []pilosa.FieldRow{{Field: "generals", RowID: 2, RowKey: "r2"}}, Count: 5, Sum: 30},
},
},
{
query: "GroupBy(Rows(generals), aggregate=Sum(field=v), having=Condition(-5<sum<27))",
expected: []pilosa.GroupCount{
{Group: []pilosa.FieldRow{{Field: "generals", RowID: 1, RowKey: "r1"}}, Count: 5, Sum: 25},
},
},
{
query: "GroupBy(Rows(generals), aggregate=Sum(field=v), having=Condition(count>5))",
expected: []pilosa.GroupCount{},
},
} }
for i, tst := range tests { for i, tst := range tests {
@ -3563,21 +3841,90 @@ func TestExecutor_Execute_Rows_Keys(t *testing.T) {
t.Run(fmt.Sprintf("#%d_%s", i, test.q), func(t *testing.T) { t.Run(fmt.Sprintf("#%d_%s", i, test.q), func(t *testing.T) {
if res, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: test.q}); err != nil { if res, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: test.q}); err != nil {
t.Fatal(err) t.Fatal(err)
} else if rows := res.Results[0].(pilosa.RowIdentifiers); !reflect.DeepEqual( } else {
rows, pilosa.RowIdentifiers{Keys: test.exp}) { rows := res.Results[0].(pilosa.RowIdentifiers)
t.Fatalf("\ngot: %+v\nexp: %+v", rows, pilosa.RowIdentifiers{Keys: test.exp}) if !reflect.DeepEqual(rows.Keys, test.exp) {
t.Fatalf("\ngot: %+v\nexp: %+v", rows.Keys, test.exp)
} else if rows.Rows != nil {
t.Fatalf("\ngot: %+v\nexp: nil", rows.Rows)
}
} }
}) })
} }
} }
func TestExecutor_ForeignIndex(t *testing.T) {
c := test.MustRunCluster(t, 1)
defer c.Close()
c.CreateField(t, "parent", pilosa.IndexOptions{Keys: true}, "general")
c.CreateField(t, "child", pilosa.IndexOptions{}, "parent_id",
pilosa.OptFieldTypeInt(0, math.MaxInt64),
pilosa.OptFieldForeignIndex("parent"),
)
c.CreateField(t, "child", pilosa.IndexOptions{}, "color",
pilosa.OptFieldKeys(),
)
// Populate parent data.
c.Query(t, "parent", `
Set("one", general=1)
Set("two", general=1)
Set("three", general=1)
Set("twenty-one", general=2)
Set("twenty-two", general=2)
Set("twenty-three", general=2)
Set("one", general=3)
Set("twenty-one", general=3)
`)
// Populate child data.
c.Query(t, "child", `
Set(1, parent_id="one")
Set(2, parent_id="two")
Set(3, parent_id="one")
Set(4, parent_id="twenty-one")
`)
// Populate color data.
c.Query(t, "child", `
Set(1, color="red")
Set(2, color="blue")
Set(3, color="blue")
Set(4, color="red")
`)
distinct := c.Query(t, "child", `Distinct(index="child", field="parent_id")`).Results[0].(pilosa.SignedRow)
if !reflect.DeepEqual(distinct.Pos.Keys, []string{"one", "two", "twenty-one"}) {
t.Fatalf("unexpected keys: %v", distinct.Pos.Keys)
}
eq := c.Query(t, "child", `Row(parent_id=="one")`).Results[0].(*pilosa.Row)
if !reflect.DeepEqual(eq.Columns(), []uint64{1, 3}) {
t.Fatalf("unexpected columns: %v", eq.Columns())
}
neq := c.Query(t, "child", `Row(parent_id!="one")`).Results[0].(*pilosa.Row)
if !reflect.DeepEqual(neq.Columns(), []uint64{2, 4}) {
t.Fatalf("unexpected columns: %v", neq.Columns())
}
join := c.Query(t, "parent", `Intersect(Row(general=3), Distinct(Row(color="blue"), index="child", field="parent_id"))`).Results[0].(*pilosa.Row)
if !reflect.DeepEqual(join.Keys, []string{"one"}) {
t.Fatalf("unexpected keys: %v", join.Keys)
}
}
func TestExecutor_Execute_GroupBy(t *testing.T) { func TestExecutor_Execute_GroupBy(t *testing.T) {
groupByTest := func(t *testing.T, clusterSize int) { groupByTest := func(t *testing.T, clusterSize int) {
c := test.MustRunCluster(t, 1) c := test.MustRunCluster(t, 1)
defer c.Close() defer c.Close()
c.CreateField(t, "i", pilosa.IndexOptions{}, "general") c.CreateField(t, "i", pilosa.IndexOptions{}, "general")
c.CreateField(t, "i", pilosa.IndexOptions{}, "sub") c.CreateField(t, "i", pilosa.IndexOptions{}, "sub")
c.CreateField(t, "i", pilosa.IndexOptions{}, "v", pilosa.OptFieldTypeInt(0, 1000))
c.ImportBits(t, "i", "general", [][2]uint64{ c.ImportBits(t, "i", "general", [][2]uint64{
{10, 0}, {10, 0},
{10, 1}, {10, 1},
@ -3596,6 +3943,11 @@ func TestExecutor_Execute_GroupBy(t *testing.T) {
{110, 2}, {110, 2},
{110, 0}, {110, 0},
}) })
if _, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(0, v=10)`}); err != nil {
t.Fatal(err)
} else if _, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `Set(1, v=100)`}); err != nil {
t.Fatal(err)
}
t.Run("No Field List Arguments", func(t *testing.T) { t.Run("No Field List Arguments", func(t *testing.T) {
if _, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `GroupBy()`}); err != nil { if _, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `GroupBy()`}); err != nil {
@ -3649,6 +4001,16 @@ func TestExecutor_Execute_GroupBy(t *testing.T) {
test.CheckGroupBy(t, expected, results) test.CheckGroupBy(t, expected, results)
}) })
t.Run("Aggregate", func(t *testing.T) {
expected := []pilosa.GroupCount{
{Group: []pilosa.FieldRow{{Field: "general", RowID: 10}, {Field: "sub", RowID: 100}}, Count: 2, Sum: 110},
{Group: []pilosa.FieldRow{{Field: "general", RowID: 10}, {Field: "sub", RowID: 110}}, Count: 1, Sum: 10},
}
results := c.Query(t, "i", `GroupBy(Rows(general), Rows(sub), aggregate=Sum(field=v))`).Results[0].([]pilosa.GroupCount)
test.CheckGroupBy(t, expected, results)
})
t.Run("check field offset no limit", func(t *testing.T) { t.Run("check field offset no limit", func(t *testing.T) {
expected := []pilosa.GroupCount{ expected := []pilosa.GroupCount{
{Group: []pilosa.FieldRow{{Field: "general", RowID: 11}}, Count: 2}, {Group: []pilosa.FieldRow{{Field: "general", RowID: 11}}, Count: 2},
@ -4090,3 +4452,186 @@ func TestExecutor_Execute_Shift(t *testing.T) {
} }
}) })
} }
func TestExecutor_Execute_IncludesColumn(t *testing.T) {
t.Run("results-ids", func(t *testing.T) {
c := test.MustRunCluster(t, 1)
defer c.Close()
hldr := test.Holder{Holder: c[0].Server.Holder()}
hldr.SetBit("i", "general", 10, 1)
hldr.SetBit("i", "general", 10, ShardWidth)
hldr.SetBit("i", "general", 10, 2*ShardWidth)
for i, tt := range []struct {
col uint64
expIncluded bool
}{
{1, true},
{2, false},
{ShardWidth, true},
{ShardWidth + 1, false},
{2 * ShardWidth, true},
{(2 * ShardWidth) + 1, false},
} {
t.Run(fmt.Sprint(i), func(t *testing.T) {
if res, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: fmt.Sprintf("IncludesColumn(Row(general=10), column=%d)", tt.col)}); err != nil {
t.Fatal(err)
} else if tt.expIncluded && !res.Results[0].(bool) {
t.Fatalf("expected to find column: %d", tt.col)
} else if !tt.expIncluded && res.Results[0].(bool) {
t.Fatalf("did not expect to find column: %d", tt.col)
}
})
}
})
t.Run("results-keys", func(t *testing.T) {
c := test.MustRunCluster(t, 1)
defer c.Close()
cmd := c[0]
hldr := test.Holder{Holder: c[0].Server.Holder()}
index := hldr.MustCreateIndexIfNotExists("i", pilosa.IndexOptions{Keys: true})
if _, err := index.CreateField("general", pilosa.OptFieldTypeDefault(), pilosa.OptFieldKeys()); err != nil {
t.Fatal(err)
}
if _, err := cmd.API.Query(
context.Background(),
&pilosa.QueryRequest{
Index: "i",
Query: `Set("one", general="ten") Set("eleven", general="ten") Set("twentyone", general="ten")`,
}); err != nil {
t.Fatal(err)
}
for i, tt := range []struct {
col string
expIncluded bool
}{
{"one", true},
{"two", false},
{"eleven", true},
{"twelve", false},
{"twentyone", true},
{"twentytwo", false},
} {
t.Run(fmt.Sprint(i), func(t *testing.T) {
if res, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: fmt.Sprintf("IncludesColumn(Row(general=ten), column=%s)", tt.col)}); err != nil {
t.Fatal(err)
} else if tt.expIncluded && !res.Results[0].(bool) {
t.Fatalf("expected to find column: %s", tt.col)
} else if !tt.expIncluded && res.Results[0].(bool) {
t.Fatalf("did not expect to find column: %s", tt.col)
}
})
}
})
t.Run("errors", func(t *testing.T) {
c := test.MustRunCluster(t, 1)
defer c.Close()
hldr := test.Holder{Holder: c[0].Server.Holder()}
hldr.SetBit("i", "general", 10, 1)
t.Run("no column", func(t *testing.T) {
expErr := "IncludesColumn call must specify a column"
if _, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `IncludesColumn(Row(general=10))`}); err == nil {
t.Fatalf("expected to get an error")
} else if !strings.Contains(err.Error(), expErr) {
t.Fatalf("expected error: %s, but got: %s", expErr, err.Error())
}
})
t.Run("no row query", func(t *testing.T) {
expErr := "IncludesColumn call must specify a row query"
if _, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `IncludesColumn(column=1)`}); err == nil {
t.Fatalf("expected to get an error")
} else if !strings.Contains(err.Error(), expErr) {
t.Fatalf("expected error: %s, but got: %s", expErr, err.Error())
}
})
})
}
func TestExecutor_Execute_MinMaxCountEqual(t *testing.T) {
c := test.MustRunCluster(t, 1)
defer c.Close()
hldr := test.Holder{Holder: c[0].Server.Holder()}
idx, err := hldr.CreateIndex("i", pilosa.IndexOptions{})
if err != nil {
t.Fatal(err)
}
if _, err := idx.CreateField("x", pilosa.OptFieldTypeDefault()); err != nil {
t.Fatal(err)
}
if _, err := idx.CreateField("f", pilosa.OptFieldTypeInt(-1100, 1000)); err != nil {
t.Fatal(err)
}
if _, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: `
Set(0, f=3)
Set(1, f=3)
Set(2, f=4)
Set(3, f=5)
Set(4, f=5)
Set(` + strconv.Itoa(ShardWidth+1) + `, f=3)
Set(` + strconv.Itoa(ShardWidth+2) + `, f=5)
Set(` + strconv.Itoa(ShardWidth+3) + `, f=5)
Set(` + strconv.Itoa(ShardWidth+4) + `, f=5)
Set(` + strconv.Itoa(ShardWidth+5) + `, f=4)
Set(` + strconv.Itoa(2*ShardWidth+1) + `, f=3)
Set(0, x=3)
Set(1, x=3)
`}); err != nil {
t.Fatal(err)
}
t.Run("Min", func(t *testing.T) {
tests := []struct {
filter string
exp int64
cnt int64
}{
{filter: ``, exp: 3, cnt: 4},
{filter: `Row(x=3)`, exp: 3, cnt: 2},
}
for i, tt := range tests {
var pql string
if tt.filter == "" {
pql = `Min(field=f)`
} else {
pql = fmt.Sprintf(`Min(%s, field=f)`, tt.filter)
}
if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: pql}); err != nil {
t.Fatal(err)
} else if !reflect.DeepEqual(result.Results[0], pilosa.ValCount{Val: tt.exp, Count: tt.cnt}) {
t.Fatalf("unexpected result, test %d: %s", i, spew.Sdump(result))
}
}
})
t.Run("Max", func(t *testing.T) {
tests := []struct {
filter string
exp int64
cnt int64
}{
{filter: ``, exp: 5, cnt: 5},
}
for i, tt := range tests {
var pql string
if tt.filter == "" {
pql = `Max(field=f)`
} else {
pql = fmt.Sprintf(`Max(%s, field=f)`, tt.filter)
}
if result, err := c[0].API.Query(context.Background(), &pilosa.QueryRequest{Index: "i", Query: pql}); err != nil {
t.Fatal(err)
} else if !reflect.DeepEqual(result.Results[0], pilosa.ValCount{Val: tt.exp, Count: tt.cnt}) {
t.Fatalf("unexpected result, test %d: %s", i, spew.Sdump(result))
}
}
})
}

96
extension.go Normal file
View file

@ -0,0 +1,96 @@
// Copyright 2019 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package pilosa
import (
"fmt"
"github.com/molecula/ext"
"github.com/pilosa/pilosa/v2/roaring"
)
// WrapBitmap yields an extension-Bitmap from a roaring Bitmap.
func WrapBitmap(bm *roaring.Bitmap) ext.Bitmap {
return wrappedBitmap{bm}
}
// wrappedBitmap is a very shallow glue shim to convert a roaring Bitmap to
// an extension Bitmap.
type wrappedBitmap struct{ *roaring.Bitmap }
// UnwrapBitmap converts an extension-bitmap to its underlying roaring Bitmap.
func UnwrapBitmap(bm ext.Bitmap) *roaring.Bitmap {
if inner, ok := bm.(wrappedBitmap); ok {
if inner.Bitmap != nil {
return inner.Bitmap
}
return roaring.NewFileBitmap()
}
return roaring.NewFileBitmap()
}
func (b wrappedBitmap) Intersect(other ext.Bitmap) ext.Bitmap {
return wrappedBitmap{b.Bitmap.Intersect(other.(wrappedBitmap).Bitmap)}
}
func (b wrappedBitmap) Union(other ext.Bitmap) ext.Bitmap {
return wrappedBitmap{b.Bitmap.Union(other.(wrappedBitmap).Bitmap)}
}
func (b wrappedBitmap) IntersectionCount(other ext.Bitmap) uint64 {
return b.Bitmap.IntersectionCount(other.(wrappedBitmap).Bitmap)
}
func (b wrappedBitmap) Difference(other ext.Bitmap) ext.Bitmap {
return wrappedBitmap{b.Bitmap.Difference(other.(wrappedBitmap).Bitmap)}
}
func (b wrappedBitmap) Xor(other ext.Bitmap) ext.Bitmap {
return wrappedBitmap{b.Bitmap.Xor(other.(wrappedBitmap).Bitmap)}
}
func (b wrappedBitmap) Shift(n int) (ext.Bitmap, error) {
shifted, err := b.Bitmap.Shift(n)
return wrappedBitmap{shifted}, err
}
func (b wrappedBitmap) Flip(start, last uint64) ext.Bitmap {
return wrappedBitmap{b.Bitmap.Flip(start, last)}
}
func (b wrappedBitmap) New() ext.Bitmap {
return WrapBitmap(roaring.NewFileBitmap())
}
// ContainerBits tries to get one container's worth of bits.
func (b wrappedBitmap) ContainerBits(offset uint64, target []uint64) (out []uint64) {
// it's an error to call this with a non-container-aligned offset
if offset&0xFFFF != 0 {
return nil
}
if b.Bitmap == nil {
fmt.Printf("ContainerBits on bitmap with no contents\n")
return nil
}
if b.Bitmap.Containers == nil {
fmt.Printf("ContainerBits on bitmap with nil Containers\n")
return nil
}
c := b.Bitmap.Containers.Get(offset >> 16)
if c == nil {
return nil
}
return c.AsBitmap(target)
}

21
extensions/distinct.go Normal file
View file

@ -0,0 +1,21 @@
// Copyright 2019 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
// +build plugindistinct
package extensions
import (
_ "github.com/molecula/extensions/distinct"
)

18
extensions/dummy.go Normal file
View file

@ -0,0 +1,18 @@
// Copyright 2019 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
// This package contains only things which are conditional on build
// tags.
package extensions

390
field.go
View file

@ -20,6 +20,7 @@ import (
"encoding/json" "encoding/json"
"fmt" "fmt"
"io/ioutil" "io/ioutil"
"math"
"os" "os"
"path/filepath" "path/filepath"
"sort" "sort"
@ -59,6 +60,7 @@ const (
FieldTypeTime = "time" FieldTypeTime = "time"
FieldTypeMutex = "mutex" FieldTypeMutex = "mutex"
FieldTypeBool = "bool" FieldTypeBool = "bool"
FieldTypeDecimal = "decimal"
) )
// Field represents a container for views. // Field represents a container for views.
@ -82,6 +84,15 @@ type Field struct {
// Field options. // Field options.
options FieldOptions options FieldOptions
// finalOptions is used with a final call to applyOptions.
// The initial call to applyOptions is made with options
// loaded from the meta file on disk (in the case when
// a field is being re-opened). If the field creator calls
// setOptions before calling Open(), then those options
// will be held in finalOptions, and applied instead of
// those from the meta file.
finalOptions *FieldOptions
bsiGroups []*bsiGroup bsiGroups []*bsiGroup
// Shards with data on any node in the cluster, according to this node. // Shards with data on any node in the cluster, according to this node.
@ -89,10 +100,18 @@ type Field struct {
logger logger.Logger logger logger.Logger
snapshotQueue chan *fragment snapshotQueue snapshotQueue
// Instantiates new translation store on open. // Instantiates new translation store on open.
OpenTranslateStore OpenTranslateStoreFunc OpenTranslateStore OpenTranslateStoreFunc
// Used for looking up a foreign index.
holder *Holder
// Stores whether or not the field has keys enabled.
// This is most helpful for cases where the keys are
// based on a foreign index; this prevents having to
// call holder.index.Keys() every time.
usesKeys bool
} }
// FieldOption is a functional option type for pilosa.fieldOptions. // FieldOption is a functional option type for pilosa.fieldOptions.
@ -107,6 +126,17 @@ func OptFieldKeys() FieldOption {
} }
} }
// OptFieldForeignIndex marks this field as a foreign key to another
// index. That is, the values of this field should be interpreted as
// referencing records (Pilosa columns) in another index. TODO explain
// where/how this is used by Pilosa.
func OptFieldForeignIndex(index string) FieldOption {
return func(fo *FieldOptions) error {
fo.ForeignIndex = index
return nil
}
}
// OptFieldTypeDefault is a functional option on FieldOptions // OptFieldTypeDefault is a functional option on FieldOptions
// used to set the field type and cache setting to the default values. // used to set the field type and cache setting to the default values.
func OptFieldTypeDefault() FieldOption { func OptFieldTypeDefault() FieldOption {
@ -155,6 +185,32 @@ func OptFieldTypeInt(min, max int64) FieldOption {
} }
} }
func OptFieldTypeDecimal(scale int64, minmax ...int64) FieldOption {
return func(fo *FieldOptions) error {
if fo.Type != "" {
return errors.Errorf("can't set field type to 'decimal', already set to: %s", fo.Type)
}
fo.Min = math.MinInt64
fo.Max = math.MaxInt64
if len(minmax) == 2 {
min, max := minmax[0], minmax[1]
if min > max {
return errors.Errorf("decimal field min cannot be greater than max, got %d, %d", min, max)
}
fo.Min = min
fo.Max = max
} else if len(minmax) > 2 {
return errors.Errorf("unknown extra parameters beyond min and max: %v", minmax)
} else if len(minmax) == 1 {
fo.Min = minmax[0]
}
fo.Type = FieldTypeDecimal
fo.Base = bsiBase(fo.Min, fo.Max)
fo.Scale = scale
return nil
}
}
// OptFieldTypeTime is a functional option on FieldOptions // OptFieldTypeTime is a functional option on FieldOptions
// used to specify the field as being type `time` and to // used to specify the field as being type `time` and to
// provide any respective configuration values. // provide any respective configuration values.
@ -203,6 +259,11 @@ func OptFieldTypeBool() FieldOption {
} }
// NewField returns a new instance of field. // NewField returns a new instance of field.
// NOTE: This function is only used in tests, which is why
// it only takes a single `FieldOption` (the assumption being
// that it's of the type `OptFieldType*`). This means
// this function couldn't be used to set, for example,
// `FieldOptions.Keys`.
func NewField(path, index, name string, opts FieldOption) (*Field, error) { func NewField(path, index, name string, opts FieldOption) (*Field, error) {
err := validateName(name) err := validateName(name)
if err != nil { if err != nil {
@ -233,7 +294,7 @@ func newField(path, index, name string, opts FieldOption) (*Field, error) {
broadcaster: NopBroadcaster, broadcaster: NopBroadcaster,
Stats: stats.NopStatsClient, Stats: stats.NopStatsClient,
options: applyDefaultOptions(fo), options: *applyDefaultOptions(&fo),
remoteAvailableShards: roaring.NewBitmap(), remoteAvailableShards: roaring.NewBitmap(),
@ -288,18 +349,30 @@ func (f *Field) mergeRemoteAvailableShards(b *roaring.Bitmap) {
// loadAvailableShards reads remoteAvailableShards data for the field, if any. // loadAvailableShards reads remoteAvailableShards data for the field, if any.
func (f *Field) loadAvailableShards() error { func (f *Field) loadAvailableShards() error {
bm := roaring.NewBitmap()
// Read data from meta file. // Read data from meta file.
path := filepath.Join(f.path, ".available.shards") path := filepath.Join(f.path, ".available.shards")
buf, err := ioutil.ReadFile(path) buf, err := ioutil.ReadFile(path)
// doesn't exist: this is fine
if os.IsNotExist(err) { if os.IsNotExist(err) {
return nil return nil
} else if err != nil {
return errors.Wrap(err, "reading available shards")
} else {
if err := bm.UnmarshalBinary(buf); err != nil {
return errors.Wrap(err, "unmarshaling")
} }
// some other problem:
if err != nil {
f.logger.Printf("available shards file present but unreadable, discarding: %v", err)
err = os.Remove(path)
if err != nil {
return errors.Wrap(err, "deleting corrupt available shards list")
}
return nil
}
bm := roaring.NewBitmap()
if err = bm.UnmarshalBinary(buf); err != nil {
f.logger.Printf("available shards file corrupt, discarding: %v", err)
err = os.Remove(path)
if err != nil {
return errors.Wrap(err, "deleting corrupt available shards list")
}
return nil
} }
// Merge bitmap from file into field. // Merge bitmap from file into field.
f.mergeRemoteAvailableShards(bm) f.mergeRemoteAvailableShards(bm)
@ -418,7 +491,13 @@ func (f *Field) Open() error {
return errors.Wrap(err, "loading available shards") return errors.Wrap(err, "loading available shards")
} }
// Apply the field options loaded from meta. // If options were provided using setOptions(), then
// use those instead of the options from the meta file.
if f.finalOptions != nil {
f.options = *f.finalOptions
}
// Apply the field options loaded from meta (or set via setOptions()).
f.logger.Debugf("apply options for index/field: %s/%s", f.index, f.name) f.logger.Debugf("apply options for index/field: %s/%s", f.index, f.name)
if err := f.applyOptions(f.options); err != nil { if err := f.applyOptions(f.options); err != nil {
return errors.Wrap(err, "applying options") return errors.Wrap(err, "applying options")
@ -434,9 +513,16 @@ func (f *Field) Open() error {
return errors.Wrap(err, "opening attrstore") return errors.Wrap(err, "opening attrstore")
} }
// Instantiate & open translation store. // If the field has a foreign index, and that index uses keys,
if f.translateStore, err = f.OpenTranslateStore(filepath.Join(f.path, "keys"), f.index, f.name); err != nil { // then use that index's translateStore instead.
return errors.Wrap(err, "opening translate store") if f.options.ForeignIndex != "" {
if err := f.holder.checkForeignIndex(f); err != nil {
return errors.Wrap(err, "checking foreign index")
}
} else {
if err := f.applyTranslateStore(); err != nil {
return errors.Wrap(err, "applying translate store")
}
} }
return nil return nil
@ -449,6 +535,35 @@ func (f *Field) Open() error {
return nil return nil
} }
// applyTranslateStore opens the configured translate store.
func (f *Field) applyTranslateStore() error {
// Instantiate & open translation store.
var err error
f.translateStore, err = f.OpenTranslateStore(filepath.Join(f.path, "keys"), f.index, f.name)
if err != nil {
return errors.Wrap(err, "opening translate store")
}
f.usesKeys = f.options.Keys
return nil
}
// applyForeignIndex sets the field's translateStore
// to that of a foreign index in the case where the
// foreign index uses keys. If the foreign index does
// not use keys, it falls back to applying the field's
// default translate store.
func (f *Field) applyForeignIndex() error {
foreignIndex := f.holder.Index(f.options.ForeignIndex)
if foreignIndex == nil {
return errors.Wrapf(ErrForeignIndexNotFound, "%s", f.options.ForeignIndex)
} else if foreignIndex.Keys() {
f.usesKeys = true
f.translateStore = foreignIndex.translateStore
return nil
}
return f.applyTranslateStore()
}
var fieldQueue = make(chan struct{}, 16) var fieldQueue = make(chan struct{}, 16)
// openViews opens and initializes the views inside the field. // openViews opens and initializes the views inside the field.
@ -550,10 +665,12 @@ func (f *Field) loadMeta() error {
f.options.Min = pb.Min f.options.Min = pb.Min
f.options.Max = pb.Max f.options.Max = pb.Max
f.options.Base = pb.Base f.options.Base = pb.Base
f.options.Scale = pb.Scale
f.options.BitDepth = uint(pb.BitDepth) f.options.BitDepth = uint(pb.BitDepth)
f.options.TimeQuantum = TimeQuantum(pb.TimeQuantum) f.options.TimeQuantum = TimeQuantum(pb.TimeQuantum)
f.options.Keys = pb.Keys f.options.Keys = pb.Keys
f.options.NoStandardView = pb.NoStandardView f.options.NoStandardView = pb.NoStandardView
f.options.ForeignIndex = pb.ForeignIndex
return nil return nil
} }
@ -584,6 +701,11 @@ func (f *Field) saveMeta() error {
return nil return nil
} }
// setOptions saves options for final application during Open().
func (f *Field) setOptions(opts *FieldOptions) {
f.finalOptions = applyDefaultOptions(opts)
}
// applyOptions configures the field based on opt. // applyOptions configures the field based on opt.
func (f *Field) applyOptions(opt FieldOptions) error { func (f *Field) applyOptions(opt FieldOptions) error {
switch opt.Type { switch opt.Type {
@ -596,29 +718,30 @@ func (f *Field) applyOptions(opt FieldOptions) error {
if opt.CacheType != "" { if opt.CacheType != "" {
f.options.CacheType = opt.CacheType f.options.CacheType = opt.CacheType
} }
if opt.CacheSize != 0 {
if opt.CacheType == CacheTypeNone { if opt.CacheType == CacheTypeNone {
f.options.CacheSize = 0 f.options.CacheSize = 0
} else { } else if opt.CacheSize != 0 {
f.options.CacheSize = opt.CacheSize f.options.CacheSize = opt.CacheSize
} }
}
f.options.Min = 0 f.options.Min = 0
f.options.Max = 0 f.options.Max = 0
f.options.Base = 0 f.options.Base = 0
f.options.BitDepth = 0 f.options.BitDepth = 0
f.options.TimeQuantum = "" f.options.TimeQuantum = ""
f.options.Keys = opt.Keys f.options.Keys = opt.Keys
case FieldTypeInt: f.options.ForeignIndex = ""
case FieldTypeInt, FieldTypeDecimal:
f.options.Type = opt.Type f.options.Type = opt.Type
f.options.CacheType = CacheTypeNone f.options.CacheType = CacheTypeNone
f.options.CacheSize = 0 f.options.CacheSize = 0
f.options.Min = opt.Min f.options.Min = opt.Min
f.options.Max = opt.Max f.options.Max = opt.Max
f.options.Base = opt.Base f.options.Base = opt.Base
f.options.Scale = opt.Scale
f.options.BitDepth = opt.BitDepth f.options.BitDepth = opt.BitDepth
f.options.TimeQuantum = "" f.options.TimeQuantum = ""
f.options.Keys = opt.Keys f.options.Keys = opt.Keys
f.options.ForeignIndex = opt.ForeignIndex
// Create new bsiGroup. // Create new bsiGroup.
bsig := &bsiGroup{ bsig := &bsiGroup{
@ -627,6 +750,7 @@ func (f *Field) applyOptions(opt FieldOptions) error {
Min: opt.Min, Min: opt.Min,
Max: opt.Max, Max: opt.Max,
Base: opt.Base, Base: opt.Base,
Scale: opt.Scale,
BitDepth: opt.BitDepth, BitDepth: opt.BitDepth,
} }
// Validate bsiGroup. // Validate bsiGroup.
@ -651,6 +775,7 @@ func (f *Field) applyOptions(opt FieldOptions) error {
f.Close() f.Close()
return errors.Wrap(err, "setting time quantum") return errors.Wrap(err, "setting time quantum")
} }
f.options.ForeignIndex = ""
case FieldTypeBool: case FieldTypeBool:
f.options.Type = FieldTypeBool f.options.Type = FieldTypeBool
f.options.CacheType = CacheTypeNone f.options.CacheType = CacheTypeNone
@ -661,6 +786,7 @@ func (f *Field) applyOptions(opt FieldOptions) error {
f.options.BitDepth = 0 f.options.BitDepth = 0
f.options.TimeQuantum = "" f.options.TimeQuantum = ""
f.options.Keys = false f.options.Keys = false
f.options.ForeignIndex = ""
default: default:
return errors.New("invalid field type") return errors.New("invalid field type")
} }
@ -695,11 +821,11 @@ func (f *Field) Close() error {
return nil return nil
} }
// keys returns true if the field uses string keys. // Keys returns true if the field uses string keys.
func (f *Field) keys() bool { func (f *Field) Keys() bool {
f.mu.RLock() f.mu.RLock()
defer f.mu.RUnlock() defer f.mu.RUnlock()
return f.options.Keys return f.usesKeys
} }
// bsiGroup returns a bsiGroup by name. // bsiGroup returns a bsiGroup by name.
@ -883,7 +1009,9 @@ func (f *Field) newView(path, name string) *view {
view.rowAttrStore = f.rowAttrStore view.rowAttrStore = f.rowAttrStore
view.stats = f.Stats view.stats = f.Stats
view.broadcaster = f.broadcaster view.broadcaster = f.broadcaster
if f.snapshotQueue != nil {
view.snapshotQueue = f.snapshotQueue view.snapshotQueue = f.snapshotQueue
}
return view return view
} }
@ -1051,6 +1179,36 @@ func (f *Field) allTimeViewsSortedByQuantum() (me []*view) {
return me return me
} }
// StringValue reads an integer field value for a column, and converts
// it to a string based on a foreign index string key.
func (f *Field) StringValue(columnID uint64) (value string, exists bool, err error) {
bsig := f.bsiGroup(f.name)
if bsig == nil {
return value, false, ErrBSIGroupNotFound
}
val, exists, err := f.Value(columnID)
if exists {
value, err = f.translateStore.TranslateID(uint64(val))
}
return value, exists, err
}
// FloatValue reads an integer field value for a column, and converts
// it to a float based on the configured scale.
func (f *Field) FloatValue(columnID uint64) (value float64, exists bool, err error) {
bsig := f.bsiGroup(f.name)
if bsig == nil {
return 0, false, ErrBSIGroupNotFound
}
val, exists, err := f.Value(columnID)
if exists {
value = float64(val) / math.Pow10(int(bsig.Scale))
}
return value, exists, err
}
// Value reads a field value for a column. // Value reads a field value for a column.
func (f *Field) Value(columnID uint64) (value int64, exists bool, err error) { func (f *Field) Value(columnID uint64) (value int64, exists bool, err error) {
bsig := f.bsiGroup(f.name) bsig := f.bsiGroup(f.name)
@ -1073,6 +1231,18 @@ func (f *Field) Value(columnID uint64) (value int64, exists bool, err error) {
return int64(v) + bsig.Base, true, nil return int64(v) + bsig.Base, true, nil
} }
// SetFloatValue takes a floating point value, and converts it to an
// integer based on the field's configured scale, before setting that
// integer via SetValue.
func (f *Field) SetFloatValue(columnID uint64, value float64) (changed bool, err error) {
bsig := f.bsiGroup(f.name)
if bsig == nil {
return false, ErrBSIGroupNotFound
}
val := int64(float64(value) * math.Pow10(int(bsig.Scale)))
return f.SetValue(columnID, val)
}
// SetValue sets a field value for a column. // SetValue sets a field value for a column.
func (f *Field) SetValue(columnID uint64, value int64) (changed bool, err error) { func (f *Field) SetValue(columnID uint64, value int64) (changed bool, err error) {
// Fetch bsiGroup & validate min/max. // Fetch bsiGroup & validate min/max.
@ -1114,10 +1284,46 @@ func (f *Field) SetValue(columnID uint64, value int64) (changed bool, err error)
if err != nil { if err != nil {
return false, errors.Wrap(err, "creating view") return false, errors.Wrap(err, "creating view")
} }
return view.setValue(columnID, bsig.BitDepth, baseValue) return view.setValue(columnID, bsig.BitDepth, baseValue)
} }
// ClearValue removes a field value for a column.
func (f *Field) ClearValue(columnID uint64) (changed bool, err error) {
bsig := f.bsiGroup(f.name)
if bsig == nil {
return false, ErrBSIGroupNotFound
}
// Fetch target view.
view := f.view(viewBSIGroupPrefix + f.name)
if view == nil {
return false, nil
}
value, exists, err := view.value(columnID, bsig.BitDepth)
if err != nil {
return false, err
}
if exists {
return view.clearValue(columnID, bsig.BitDepth, value)
}
return false, nil
}
// FloatSum performs a Sum query and converts the result to a float
// based on the field's configured scale.
func (f *Field) FloatSum(filter *Row, name string) (sum float64, count int64, err error) {
bsig := f.bsiGroup(f.name)
if bsig == nil {
return 0, 0, ErrBSIGroupNotFound
}
sumI, count, err := f.Sum(filter, name)
if err == nil {
sum = float64(sumI) / math.Pow10(int(bsig.Scale))
}
return sum, count, err
}
// Sum returns the sum and count for a field. // Sum returns the sum and count for a field.
// An optional filtering row can be provided. // An optional filtering row can be provided.
func (f *Field) Sum(filter *Row, name string) (sum, count int64, err error) { func (f *Field) Sum(filter *Row, name string) (sum, count int64, err error) {
@ -1138,6 +1344,21 @@ func (f *Field) Sum(filter *Row, name string) (sum, count int64, err error) {
return int64(vsum) + (int64(vcount) * bsig.Base), int64(vcount), nil return int64(vsum) + (int64(vcount) * bsig.Base), int64(vcount), nil
} }
// FloatMin performs a Min query and converts the result to a float
// based on the field's configured scale.
func (f *Field) FloatMin(filter *Row, name string) (min float64, count int64, err error) {
bsig := f.bsiGroup(f.name)
if bsig == nil {
return 0, 0, ErrBSIGroupNotFound
}
minI, count, err := f.Min(filter, name)
if err == nil {
min = float64(minI) / math.Pow10(int(bsig.Scale))
}
return min, count, err
}
// Min returns the min for a field. // Min returns the min for a field.
// An optional filtering row can be provided. // An optional filtering row can be provided.
func (f *Field) Min(filter *Row, name string) (min, count int64, err error) { func (f *Field) Min(filter *Row, name string) (min, count int64, err error) {
@ -1158,6 +1379,21 @@ func (f *Field) Min(filter *Row, name string) (min, count int64, err error) {
return int64(vmin) + bsig.Base, int64(vcount), nil return int64(vmin) + bsig.Base, int64(vcount), nil
} }
// FloatMax performs a max query and converts the result to a float
// based on the field's configured scale.
func (f *Field) FloatMax(filter *Row, name string) (max float64, count int64, err error) {
bsig := f.bsiGroup(f.name)
if bsig == nil {
return 0, 0, ErrBSIGroupNotFound
}
maxI, count, err := f.Max(filter, name)
if err == nil {
max = float64(maxI) / math.Pow10(int(bsig.Scale))
}
return max, count, err
}
// Max returns the max for a field. // Max returns the max for a field.
// An optional filtering row can be provided. // An optional filtering row can be provided.
func (f *Field) Max(filter *Row, name string) (max, count int64, err error) { func (f *Field) Max(filter *Row, name string) (max, count int64, err error) {
@ -1283,6 +1519,21 @@ func (f *Field) Import(rowIDs, columnIDs []uint64, timestamps []*time.Time, opts
return nil return nil
} }
func (f *Field) importFloatValue(columnIDs []uint64, values []float64, options *ImportOptions) error {
// convert values to int64 values based on scale
ivalues := make([]int64, len(values))
bsig := f.bsiGroup(f.name)
if bsig == nil {
return errors.Wrap(ErrBSIGroupNotFound, f.name)
}
mult := math.Pow10(int(bsig.Scale))
for i, fval := range values {
ivalues[i] = int64(fval * mult)
}
// then call importValue
return f.importValue(columnIDs, ivalues, options)
}
// importValue bulk imports range-encoded value data. // importValue bulk imports range-encoded value data.
func (f *Field) importValue(columnIDs []uint64, values []int64, options *ImportOptions) error { func (f *Field) importValue(columnIDs []uint64, values []int64, options *ImportOptions) error {
viewName := viewBSIGroupPrefix + f.name viewName := viewBSIGroupPrefix + f.name
@ -1292,34 +1543,12 @@ func (f *Field) importValue(columnIDs []uint64, values []int64, options *ImportO
return errors.Wrap(ErrBSIGroupNotFound, f.name) return errors.Wrap(ErrBSIGroupNotFound, f.name)
} }
// Find the lowest/highest values. // We want to determine the required bit depth, in case the field doesn't
// have as many bits currently as would be needed to represent these values,
// but only if the values are in-range for the field.
var min, max int64 var min, max int64
for i, value := range values { if len(values) > 0 {
if i == 0 || value < min { min, max = values[0], values[0]
min = value
}
if i == 0 || value > max {
max = value
}
}
// Determine the highest bit depth required by the min & max.
requiredDepth := bitDepthInt64(min - bsig.Base)
if v := bitDepthInt64(max - bsig.Base); v > requiredDepth {
requiredDepth = v
}
// Increase bit depth if required.
if requiredDepth > bsig.BitDepth {
if err := func() error {
f.mu.Lock()
defer f.mu.Unlock()
bsig.BitDepth = requiredDepth
f.options.BitDepth = requiredDepth
return f.saveMeta()
}(); err != nil {
return errors.Wrap(err, "increasing bsi bit depth")
}
} }
// Split import data by fragment. // Split import data by fragment.
@ -1331,6 +1560,12 @@ func (f *Field) importValue(columnIDs []uint64, values []int64, options *ImportO
} else if value < bsig.Min { } else if value < bsig.Min {
return fmt.Errorf("%v, columnID=%v, value=%v", ErrBSIGroupValueTooLow, columnID, value) return fmt.Errorf("%v, columnID=%v, value=%v", ErrBSIGroupValueTooLow, columnID, value)
} }
if value > max {
max = value
}
if value < min {
min = value
}
// Attach value to each bsiGroup view. // Attach value to each bsiGroup view.
for _, name := range []string{viewName} { for _, name := range []string{viewName} {
@ -1342,6 +1577,26 @@ func (f *Field) importValue(columnIDs []uint64, values []int64, options *ImportO
} }
} }
// Determine the highest bit depth required by the min & max.
requiredDepth := bitDepthInt64(min - bsig.Base)
if v := bitDepthInt64(max - bsig.Base); v > requiredDepth {
requiredDepth = v
}
// Increase bit depth if required.
if requiredDepth > bsig.BitDepth {
if err := func() error {
f.mu.Lock()
defer f.mu.Unlock()
bsig.BitDepth = requiredDepth
f.options.BitDepth = requiredDepth
return f.saveMeta()
}(); err != nil {
return errors.Wrap(err, "increasing bsi bit depth")
}
} else {
requiredDepth = bsig.BitDepth
}
// Import into each fragment. // Import into each fragment.
for key, data := range dataByFragment { for key, data := range dataByFragment {
// The view must already exist (i.e. we can't create it) // The view must already exist (i.e. we can't create it)
@ -1419,23 +1674,23 @@ type FieldOptions struct {
BitDepth uint `json:"bitDepth,omitempty"` BitDepth uint `json:"bitDepth,omitempty"`
Min int64 `json:"min,omitempty"` Min int64 `json:"min,omitempty"`
Max int64 `json:"max,omitempty"` Max int64 `json:"max,omitempty"`
Scale int64 `json:"scale,omitempty"`
Keys bool `json:"keys"` Keys bool `json:"keys"`
NoStandardView bool `json:"noStandardView,omitempty"` NoStandardView bool `json:"noStandardView,omitempty"`
CacheSize uint32 `json:"cacheSize,omitempty"` CacheSize uint32 `json:"cacheSize,omitempty"`
CacheType string `json:"cacheType,omitempty"` CacheType string `json:"cacheType,omitempty"`
Type string `json:"type,omitempty"` Type string `json:"type,omitempty"`
TimeQuantum TimeQuantum `json:"timeQuantum,omitempty"` TimeQuantum TimeQuantum `json:"timeQuantum,omitempty"`
ForeignIndex string `json:"foreignIndex"`
} }
// applyDefaultOptions returns a new FieldOptions object // applyDefaultOptions updates FieldOptions with the default
// with default values if o does not contain a valid type. // values if o does not contain a valid type.
func applyDefaultOptions(o FieldOptions) FieldOptions { func applyDefaultOptions(o *FieldOptions) *FieldOptions {
if o.Type == "" { if o.Type == "" {
return FieldOptions{ o.Type = DefaultFieldType
Type: DefaultFieldType, o.CacheType = DefaultCacheType
CacheType: DefaultCacheType, o.CacheSize = DefaultCacheSize
CacheSize: DefaultCacheSize,
}
} }
return o return o
} }
@ -1454,12 +1709,14 @@ func encodeFieldOptions(o *FieldOptions) *internal.FieldOptions {
CacheType: o.CacheType, CacheType: o.CacheType,
CacheSize: o.CacheSize, CacheSize: o.CacheSize,
Base: o.Base, Base: o.Base,
Scale: o.Scale,
BitDepth: uint64(o.BitDepth), BitDepth: uint64(o.BitDepth),
Min: o.Min, Min: o.Min,
Max: o.Max, Max: o.Max,
TimeQuantum: string(o.TimeQuantum), TimeQuantum: string(o.TimeQuantum),
Keys: o.Keys, Keys: o.Keys,
NoStandardView: o.NoStandardView, NoStandardView: o.NoStandardView,
ForeignIndex: o.ForeignIndex,
} }
} }
@ -1488,6 +1745,7 @@ func (o *FieldOptions) MarshalJSON() ([]byte, error) {
Min int64 `json:"min"` Min int64 `json:"min"`
Max int64 `json:"max"` Max int64 `json:"max"`
Keys bool `json:"keys"` Keys bool `json:"keys"`
ForeignIndex string `json:"foreignIndex"`
}{ }{
o.Type, o.Type,
o.Base, o.Base,
@ -1495,6 +1753,25 @@ func (o *FieldOptions) MarshalJSON() ([]byte, error) {
o.Min, o.Min,
o.Max, o.Max,
o.Keys, o.Keys,
o.ForeignIndex,
})
case FieldTypeDecimal:
return json.Marshal(struct {
Type string `json:"type"`
Base int64 `json:"base"`
Scale int64 `json:"scale"`
BitDepth uint `json:"bitDepth"`
Min int64 `json:"min"`
Max int64 `json:"max"`
Keys bool `json:"keys"`
}{
o.Type,
o.Base,
o.Scale,
o.BitDepth,
o.Min,
o.Max,
o.Keys,
}) })
case FieldTypeTime: case FieldTypeTime:
return json.Marshal(struct { return json.Marshal(struct {
@ -1563,6 +1840,7 @@ type bsiGroup struct {
Min int64 `json:"min,omitempty"` Min int64 `json:"min,omitempty"`
Max int64 `json:"max,omitempty"` Max int64 `json:"max,omitempty"`
Base int64 `json:"base,omitempty"` Base int64 `json:"base,omitempty"`
Scale int64 `json:"scale,omitempty"`
BitDepth uint `json:"bitDepth,omitempty"` BitDepth uint `json:"bitDepth,omitempty"`
} }

View file

@ -19,6 +19,7 @@ import (
"io/ioutil" "io/ioutil"
"math" "math"
"os" "os"
"path/filepath"
"reflect" "reflect"
"testing" "testing"
"time" "time"
@ -366,6 +367,62 @@ func TestField_PersistAvailableShards(t *testing.T) {
} }
func TestField_CorruptAvailableShards(t *testing.T) {
f := MustOpenField(OptFieldTypeDefault())
// bm represents remote available shards.
bm := roaring.NewBitmap(1, 2, 3)
if err := f.AddRemoteAvailableShards(bm); err != nil {
t.Fatal(err)
}
path := filepath.Join(f.path, ".available.shards")
avail, err := os.OpenFile(path, os.O_APPEND|os.O_WRONLY, 0644)
if err != nil {
t.Fatal(err)
}
n, err := avail.Write([]byte{23})
if err != nil || n != 1 {
t.Fatal(err)
}
avail.Close()
// Reload field and verify that shard data is persisted.
if err := f.Reopen(); err != nil {
t.Fatal(err)
} else if !reflect.DeepEqual(f.remoteAvailableShards.Slice(), []uint64(nil)) {
t.Fatalf("unexpected available shards (reopen). expected: %#v, but got: %#v", []uint64{}, f.remoteAvailableShards.Slice())
}
}
func TestField_TruncatedAvailableShards(t *testing.T) {
f := MustOpenField(OptFieldTypeDefault())
// bm represents remote available shards.
bm := roaring.NewBitmap(1, 2, 3)
if err := f.AddRemoteAvailableShards(bm); err != nil {
t.Fatal(err)
}
path := filepath.Join(f.path, ".available.shards")
avail, err := os.OpenFile(path, os.O_TRUNC|os.O_WRONLY, 0644)
if err != nil {
t.Fatal(err)
}
avail.Close()
// Reload field and verify that shard data is persisted.
if err := f.Reopen(); err != nil {
t.Fatal(err)
} else if !reflect.DeepEqual(f.remoteAvailableShards.Slice(), []uint64(nil)) {
t.Fatalf("unexpected available shards (reopen). expected: %#v, but got: %#v", []uint64{}, f.remoteAvailableShards.Slice())
}
}
// Ensure that persisting available shards having a smaller footprint (for example, // Ensure that persisting available shards having a smaller footprint (for example,
// when going from a bitmap to a smaller, RLE representation) succeeds. // when going from a bitmap to a smaller, RLE representation) succeeds.
func TestField_PersistAvailableShardsFootprint(t *testing.T) { func TestField_PersistAvailableShardsFootprint(t *testing.T) {
@ -438,3 +495,37 @@ func TestBSIGroup_BaseDefaultValue(t *testing.T) {
} }
} }
} }
func TestField_ApplyOptions(t *testing.T) {
for i, tt := range []struct {
opts FieldOptions
expOpts FieldOptions
}{
{
FieldOptions{
Type: FieldTypeSet,
CacheType: CacheTypeNone,
CacheSize: 0,
},
FieldOptions{
Type: FieldTypeSet,
CacheType: CacheTypeNone,
CacheSize: 0,
},
},
} {
fld := &Field{}
fld.options = *applyDefaultOptions(&FieldOptions{})
if err := fld.applyOptions(tt.opts); err != nil {
t.Fatal(err)
}
if fld.options.CacheType != tt.expOpts.CacheType {
t.Fatalf("test %d, unexpected FieldOptions.CacheType value. expected: %s, but got: %s", i, tt.expOpts.CacheType, fld.options.CacheType)
} else if fld.options.CacheSize != tt.expOpts.CacheSize {
t.Fatalf("test %d, unexpected FieldOptions.CacheSize value. expected: %d, but got: %d", i, tt.expOpts.CacheSize, fld.options.CacheSize)
}
}
}

View file

@ -158,6 +158,7 @@ func TestField_NameValidation(t *testing.T) {
"under_score", "under_score",
"abc123", "abc123",
"trailing_", "trailing_",
"charact2301234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890",
} }
invalidFieldNames := []string{ invalidFieldNames := []string{
"", "",
@ -168,7 +169,7 @@ func TestField_NameValidation(t *testing.T) {
"abc def", "abc def",
"camelCase", "camelCase",
"UPPERCASE", "UPPERCASE",
"a12345678901234567890123456789012345678901234567890123456789012345", "charact23112345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901",
} }
path, err := ioutil.TempDir("", "pilosa-field-") path, err := ioutil.TempDir("", "pilosa-field-")
@ -227,3 +228,44 @@ func TestField_AvailableShards(t *testing.T) {
t.Fatal(diff) t.Fatal(diff)
} }
} }
func TestField_ClearValue(t *testing.T) {
t.Run("OK", func(t *testing.T) {
idx := test.MustOpenIndex()
defer idx.Close()
f, err := idx.CreateField("f", pilosa.OptFieldTypeInt(math.MinInt64, math.MaxInt64))
if err != nil {
t.Fatal(err)
}
// Set value on field.
if changed, err := f.SetValue(100, 21); err != nil {
t.Fatal(err)
} else if !changed {
t.Fatal("expected change")
}
// Read value.
if value, exists, err := f.Value(100); err != nil {
t.Fatal(err)
} else if value != 21 {
t.Fatalf("unexpected value: %d", value)
} else if !exists {
t.Fatal("expected value to exist")
}
if changed, err := f.ClearValue(100); err != nil {
t.Fatal(err)
} else if !changed {
t.Fatal(err)
}
// Read value.
if _, exists, err := f.Value(100); err != nil {
t.Fatal(err)
} else if exists {
t.Fatal("expected value to not exist")
}
})
}

View file

@ -26,6 +26,7 @@ import (
"io" "io"
"io/ioutil" "io/ioutil"
"math" "math"
"math/bits"
"os" "os"
"sort" "sort"
"strings" "strings"
@ -42,7 +43,6 @@ import (
"github.com/pilosa/pilosa/v2/roaring" "github.com/pilosa/pilosa/v2/roaring"
"github.com/pilosa/pilosa/v2/shardwidth" "github.com/pilosa/pilosa/v2/shardwidth"
"github.com/pilosa/pilosa/v2/stats" "github.com/pilosa/pilosa/v2/stats"
"github.com/pilosa/pilosa/v2/syswrap"
"github.com/pilosa/pilosa/v2/tracing" "github.com/pilosa/pilosa/v2/tracing"
"github.com/pkg/errors" "github.com/pkg/errors"
) )
@ -109,19 +109,14 @@ type fragment struct {
// File-backed storage // File-backed storage
path string path string
flags byte // user-defined flags passed to roaring flags byte // user-defined flags passed to roaring
file *os.File gen generation
storage *roaring.Bitmap storage *roaring.Bitmap
storageData []byte
totalOpN int64 // total opN values
totalOps int64 // total ops (across all snapshots)
opN int // number of ops since snapshot (may be approximate for imports) opN int // number of ops since snapshot (may be approximate for imports)
ops int // number of higher-level operations, as opposed to bit changes ops int // number of higher-level operations, as opposed to bit changes
snapshotsRequested int // number of times we've requested a snapshot snapshotPending bool // set to true when requesting a snapshot, set to false after snapshot completes
snapshotsTaken int // number of actual snapshot operations
snapshotting bool // set to true when requesting a snapshot, set to false after snapshot completes
snapshotCond sync.Cond snapshotCond sync.Cond
snapshotDelays int snapshotErr error // error yielded by the last snapshot operation
snapshotDelayTime time.Duration snapshotStamp time.Time // timestamp of last snapshot
// Cache for row counts. // Cache for row counts.
CacheType string // passed in by field CacheType string // passed in by field
@ -155,7 +150,7 @@ type fragment struct {
stats stats.StatsClient stats stats.StatsClient
snapshotQueue chan *fragment snapshotQueue snapshotQueue
} }
// newFragment returns a new instance of Fragment. // newFragment returns a new instance of Fragment.
@ -174,6 +169,7 @@ func newFragment(path, index, field, view string, shard uint64, flags byte) *fra
MaxOpN: defaultFragmentMaxOpN, MaxOpN: defaultFragmentMaxOpN,
stats: stats.NopStatsClient, stats: stats.NopStatsClient,
snapshotQueue: defaultSnapshotQueue,
} }
f.snapshotCond = sync.Cond{L: &f.mu} f.snapshotCond = sync.Cond{L: &f.mu}
return f return f
@ -182,62 +178,6 @@ func newFragment(path, index, field, view string, shard uint64, flags byte) *fra
// cachePath returns the path to the fragment's cache data. // cachePath returns the path to the fragment's cache data.
func (f *fragment) cachePath() string { return f.path + cacheExt } func (f *fragment) cachePath() string { return f.path + cacheExt }
// newSnapshotQueue makes a new snapshot queue, of depth N, and spawns a
// goroutine for it.
func newSnapshotQueue(n int, w int, l logger.Logger) chan *fragment {
ch := make(chan *fragment, n)
for i := 0; i < w; i++ {
go snapshotQueueWorker(ch, l)
}
return ch
}
func snapshotQueueWorker(snapshotQueue chan *fragment, l logger.Logger) {
for f := range snapshotQueue {
err := f.protectedSnapshot(true)
if err != nil {
l.Printf("snapshot error: %v", err)
}
f.snapshotCond.Broadcast()
}
}
// enqueueSnapshot requests that the fragment be snapshotted at some point
// in the future, if this has not already been requested. Call this only when
// the mutex is held.
func (f *fragment) enqueueSnapshot() {
f.snapshotsRequested++
if f.snapshotting {
return
}
f.snapshotting = true
if f.snapshotQueue != nil {
select {
case f.snapshotQueue <- f:
default:
before := time.Now()
// wait forever, but notice that we're waiting
f.snapshotQueue <- f
f.snapshotDelays++
f.snapshotDelayTime += time.Since(before)
if f.snapshotDelays >= 10 {
f.Logger.Printf("snapshotting %s: last ten enqueue delays took %v", f.path, f.snapshotDelayTime)
f.snapshotDelays = 0
f.snapshotDelayTime = 0
}
}
} else {
// in testing, for instance, there may be no holder, thus no one
// to handle these snapshots.
err := f.snapshot()
if err != nil {
f.Logger.Printf("snapshot failed: %v", err)
}
f.snapshotting = false
f.snapshotCond.Broadcast()
}
}
// Open opens the underlying storage. // Open opens the underlying storage.
func (f *fragment) Open() error { func (f *fragment) Open() error {
f.mu.Lock() f.mu.Lock()
@ -253,6 +193,10 @@ func (f *fragment) Open() error {
// Fill cache with rows persisted to disk. // Fill cache with rows persisted to disk.
f.Logger.Debugf("open cache for index/field/view/fragment: %s/%s/%s/%d", f.index, f.field, f.view, f.shard) f.Logger.Debugf("open cache for index/field/view/fragment: %s/%s/%s/%d", f.index, f.field, f.view, f.shard)
if err := f.openCache(); err != nil { if err := f.openCache(); err != nil {
e2 := f.closeStorage()
if e2 != nil {
return errors.Wrapf(err, "closing storage: %v, after opening cache", e2)
}
return errors.Wrap(err, "opening cache") return errors.Wrap(err, "opening cache")
} }
@ -272,48 +216,103 @@ func (f *fragment) Open() error {
return nil return nil
} }
func (f *fragment) reopen() (mustClose bool, err error) { // emptyStorage is the common case for importStorage/applyStorage where they
if f.file == nil { // get no data. It tries to write the current storage to the provided file,
// Open the data file to be mmap'd and used as an ops log. // which is assumed to be the file they didn't get any data from.
f.file, mustClose, err = syswrap.OpenFile(f.path, os.O_RDWR|os.O_CREATE|os.O_APPEND, 0666) func (f *fragment) emptyStorage(file *os.File) (bool, error) {
// No data. We'll mark this for no mapping, clear any existing
// mapped containers, and set the Source to nil. We also have no
// ops.
f.opN = 0
f.ops = 0
f.storage.SetOps(0, 0)
f.storage.PreferMapping(false)
_, err := f.storage.RemapRoaringStorage(nil)
f.storage.SetSource(nil)
if err != nil { if err != nil {
return mustClose, fmt.Errorf("open file: %s", err) return false, fmt.Errorf("applying/importing storage: no data, and clearing old mapping also failed: %v", err)
} }
f.storage.OpWriter = f.file // Write the existing storage out to the file so it's
// a valid Roaring file thereafter. nothing to unmarshal.
// In the unlikely event that this happened even though we
// had significant data, we're not mapping it, but that's
// harmless even if it's not maximally efficient.
bi := bufio.NewWriter(file)
if _, err = f.storage.WriteTo(bi); err != nil {
return false, fmt.Errorf("init storage file: %s", err)
} }
return mustClose, nil bi.Flush()
return false, nil
} }
// openStorage opens the storage bitmap. Usually you also want to read in // importStorage attempts to import data from storage -- for instance,
// the storage, but in the case where we just wrote that file, such as // reading in a roaring bitmap from media.
// unprotectedWriteToFragment, we could also just... not. If we didn't func (f *fragment) importStorage(data []byte, file *os.File, newGen generation, mapped bool) (bool, error) {
// have existing storage, we probably need to unmarshal the data. If the f.storage.PreferMapping(mapped)
// file we're asked to open is empty, we probably don't. if len(data) == 0 {
// return f.emptyStorage(file)
// If we already had mapped storage previously, we want to unmap that, and }
// possibly remap it from the file, but we don't need a full unmarshal, just
// an update of mapped pointers.
//
// unmarshalData is somewhat overloaded. it tells us whether or not we
// need to actually create a bitmap from the data (if the data exists to
// do this from).
//
// usually unmarshalData is only set to false when we're in the middle of
// a snapshot, and unprotectedWriteToFragment just wrote the in-memory data
// out.
//
// If we have existing storage data, and we successfully get new data,
// we will unmap the existing storage data.
//
// This function's design is probably a problem -- it is trying to handle
// both cases where there was existing data before, and cases where we
// just wrote the data.
func (f *fragment) openStorage(unmarshalData bool) error {
oldStorageData := f.storageData
// there's a few places where we might encounter an error, but need
// to continue past it through other error checks, before returning it.
var lastError error
// UnmarshalBinary will have remapped the storage to newGen if it
// succeeded, or if it fails but the error is advisory-only. So we
// optimistically set the source here, but if there's a non-advisory
// error, we'll unmap it and then set the source to nil.
f.storage.SetSource(newGen)
if err := f.storage.UnmarshalBinary(data); err != nil {
// roaring can report advisory-only errors...
cause := errors.Cause(err)
_, ok := cause.(roaring.AdvisoryError)
if !ok {
_, e2 := f.storage.RemapRoaringStorage(nil)
f.storage.SetSource(nil)
if e2 != nil {
return false, fmt.Errorf("unmarshal storage: file=%s, err=%s, clearing old mapping also failed: %v", file.Name(), err, e2)
}
return false, fmt.Errorf("unmarshal storage: file=%s, err=%s", file.Name(), err)
}
f.Logger.Printf("warning: unmarshal storage, file=%s, err=%v", file.Name(), err)
trunc, ok := cause.(roaring.FileShouldBeTruncatedError)
if ok {
// generation code looks for a FileShouldBeTruncatedError
return false, trunc
}
}
f.ops, f.opN = f.storage.Ops()
// For now, we assume that UnmarshalBinary will have mapped at least
// one container if we told it the storage was mapped and it didn't
// error out. This might be wrong in occasional trivial cases, but
// it should be harmless.
return mapped, nil
}
// applyStorage applies storage to a fragment that may already have
// usable data. For instance, this would try to remap existing containers
// to use a new storage as backing store.
func (f *fragment) applyStorage(data []byte, file *os.File, newGen generation, mapped bool) (bool, error) {
if len(data) == 0 {
return f.emptyStorage(file)
}
// Tell storage to prefer mapping if and only if we think the data
// is mmapped and valid.
f.storage.PreferMapping(mapped)
f.storage.SetSource(newGen)
// RemapRoaringStorage will fix any mapped containers to point either
// to the provided data (if PreferMapping was called with true and
// data is provided and there's a corresponding container) or to
// allocated storage, so when it's done, there's nothing in it that
// is mapped to anything *other than* the provided data.
return f.storage.RemapRoaringStorage(data)
}
// openStorage opens the storage bitmap.
//
// This has been massively reworked recently, and now hands a lot of
// file management off to the generation object and the Done method
// of that object. Similarly, the bitmap mapping/remapping
// logic is now mostly in importStorage (reading in a bitmap) and applyStorage
// (remapping an existing bitmap to match a new backing store).
func (f *fragment) openStorage(unmarshalData bool) error {
// Create a roaring bitmap to serve as storage for the shard. // Create a roaring bitmap to serve as storage for the shard.
if f.storage == nil { if f.storage == nil {
f.storage = roaring.NewFileBitmap() f.storage = roaring.NewFileBitmap()
@ -322,139 +321,23 @@ func (f *fragment) openStorage(unmarshalData bool) error {
// unmarshal this data in order to have any. // unmarshal this data in order to have any.
unmarshalData = true unmarshalData = true
} }
// Open the data file to be mmap'd and used as an ops log. f.rowCache = &simpleCache{make(map[uint64]*Row)}
file, mustClose, err := syswrap.OpenFile(f.path, os.O_RDWR|os.O_CREATE|os.O_APPEND, 0666) var storageOp func([]byte, *os.File, generation, bool) (bool, error)
if err != nil { if unmarshalData {
return fmt.Errorf("open file: %s", err) storageOp = f.importStorage
} else {
storageOp = f.applyStorage
} }
f.file = file
if mustClose {
defer f.safeClose()
}
// Lock the underlying file.
if err := syscall.Flock(int(f.file.Fd()), syscall.LOCK_EX|syscall.LOCK_NB); err != nil {
return fmt.Errorf("flock: %s", err)
}
// data is the data we would unmarshal from, if we're unmarshalling; it might
// be obtained by calling ReadAll on a file.
//
// newStorageData is the data we should map things to. it is set only if
// mmapped; if we didn't mmap (say, we couldn't), we won't want to unmap
// the ioutil byte slice. (Theoretically, we shouldn't be using the mapped
// flag in that case...)
var data []byte
var newStorageData []byte
// If the file is empty then initialize it with an empty bitmap.
fi, err := f.file.Stat()
if err != nil {
return errors.Wrap(err, "statting file before")
} else if fi.Size() == 0 {
bi := bufio.NewWriter(f.file)
var err error var err error
if _, err = f.storage.WriteTo(bi); err != nil { f.gen, err = newGeneration(f.gen, f.path, unmarshalData, storageOp, f.Logger)
return fmt.Errorf("init storage file: %s", err) if generationDebug {
// We might have already done this anyway, if we think we
// mapped stuff, but when debugging we want to do it
// unconditionally, because the test cases otherwise won't
// exercise this code well.
f.storage.SetSource(f.gen)
} }
bi.Flush() return err
_, err = f.file.Stat()
if err != nil {
return errors.Wrap(err, "statting file after")
}
// there's nothing here, we're not going to try to unmarshal it.
unmarshalData = false
f.rowCache = &simpleCache{make(map[uint64]*Row)}
} else {
// Mmap the underlying file so it can be zero copied.
data, err = syswrap.Mmap(int(f.file.Fd()), 0, int(fi.Size()), syscall.PROT_READ, syscall.MAP_SHARED)
if err == syswrap.ErrMaxMapCountReached {
f.Logger.Debugf("maximum number of maps reached, reading file instead")
if unmarshalData {
data, err = ioutil.ReadAll(file)
if err != nil {
return errors.Wrap(err, "failure file readall")
}
}
} else if err != nil {
return errors.Wrap(err, "mmap failed")
} else {
newStorageData = data
}
}
if unmarshalData {
f.storageData = newStorageData
// We're about to either re-read the bitmap, or fail to do so
// and unconditionally unmap the existing stuff. Either way, we
// want to unmap the old storage data after we're done here, but
// we can't unmap it yet because it's still live until sometime
// later, but we can't unmap it later, because we could return
// early... this is what defer is for.
if oldStorageData != nil {
defer func() {
unmapErr := syswrap.Munmap(oldStorageData)
if unmapErr != nil {
f.Logger.Printf("unmap of old storage failed: %s", err)
}
}()
}
// set the preference for mapping based on whether the data's mmapped
f.storage.PreferMapping(newStorageData != nil)
// so we have a problem here: if this fails, it's unclear whether
// *either* or *both* of old and new storage data might be in use.
// So we call the thing that should unconditionally unmap both of them...
if err := f.storage.UnmarshalBinary(data); err != nil {
_, e2 := f.storage.RemapRoaringStorage(nil)
if e2 != nil {
return fmt.Errorf("unmarshal storage: file=%s, err=%s, clearing old mapping also failed: %v", f.file.Name(), err, e2)
}
return fmt.Errorf("unmarshal storage: file=%s, err=%s", f.file.Name(), err)
}
f.rowCache = &simpleCache{make(map[uint64]*Row)}
f.ops, f.opN = f.storage.Ops()
} else {
// we're moving to new storage, so instead of using the OpN
// derived from reading that storage, we notify the bitmap that
// OpN is now effectively zero.
f.opN = 0
f.ops = 0
f.storage.SetOps(0, 0)
// if oldStorageData is nil, this just tries to unmap any bits that
// are currently mapped. otherwise, it will point them at this
// storage (if the containers match).
var mappedAny bool
mappedAny, lastError = f.storage.RemapRoaringStorage(newStorageData)
if oldStorageData != nil {
unmapErr := syswrap.Munmap(oldStorageData)
if unmapErr != nil {
f.Logger.Printf("unmap of old storage failed: %s", err)
}
}
if mappedAny {
// Advise the kernel that the mmap is accessed randomly.
if err := madvise(newStorageData, syscall.MADV_RANDOM); err != nil {
lastError = fmt.Errorf("madvise: %s", err)
}
} else {
// if we did map data, but for some reason none of it got used
// as backing store, we can unmap it, and set the slice to nil,
// so we don't keep the now-invalid slice in f.storageData.
if newStorageData != nil {
unmapErr := syswrap.Munmap(newStorageData)
if unmapErr != nil {
lastError = fmt.Errorf("unmapping unused storage data: %s", err)
}
newStorageData = nil
}
}
f.storageData = newStorageData
}
// Attach the file to the bitmap to act as a write-ahead log.
f.storage.OpWriter = f.file
return lastError
} }
// openCache initializes the cache from row ids persisted to disk. // openCache initializes the cache from row ids persisted to disk.
@ -503,31 +386,12 @@ func (f *fragment) openCache() error {
func (f *fragment) Close() error { func (f *fragment) Close() error {
f.mu.Lock() f.mu.Lock()
defer f.mu.Unlock() defer f.mu.Unlock()
for f.snapshotting { for f.snapshotPending {
f.snapshotCond.Wait() f.snapshotCond.Wait()
} }
return f.close() return f.close()
} }
// awaitSnapshot lets us delay until the snapshot gets written, preventing tests
// from misleadingly showing amazingly fast performance because the snapshots they
// trigger haven't happened yet.
func (f *fragment) awaitSnapshot() {
f.mu.Lock()
defer f.mu.Unlock()
for f.snapshotting {
f.snapshotCond.Wait()
}
}
// unprotectedAwaitSnapshot assumes you already hold the lock, and waits for
// the snapshot fairy to come along.
func (f *fragment) unprotectedAwaitSnapshot() {
for f.snapshotting {
f.snapshotCond.Wait()
}
}
func (f *fragment) close() error { func (f *fragment) close() error {
// Flush cache if closing gracefully. // Flush cache if closing gracefully.
if err := f.flushCache(); err != nil { if err := f.flushCache(); err != nil {
@ -536,7 +400,7 @@ func (f *fragment) close() error {
} }
// Close underlying storage. // Close underlying storage.
if err := f.closeStorage(true); err != nil { if err := f.closeStorage(); err != nil {
f.Logger.Printf("fragment: error closing storage: err=%s, path=%s", err, f.path) f.Logger.Printf("fragment: error closing storage: err=%s, path=%s", err, f.path)
return errors.Wrap(err, "closing storage") return errors.Wrap(err, "closing storage")
} }
@ -547,54 +411,17 @@ func (f *fragment) close() error {
return nil return nil
} }
// safeClose is unprotected. // closeStorage marks the current generation as done. It is not necessary
func (f *fragment) safeClose() error { // to call this before openStorage.
// Flush file, unlock & close. func (f *fragment) closeStorage() error {
if f.file != nil {
if err := f.file.Sync(); err != nil {
return fmt.Errorf("sync: %s", err)
}
if err := syscall.Flock(int(f.file.Fd()), syscall.LOCK_UN); err != nil {
return fmt.Errorf("unlock: %s", err)
}
if err := syswrap.CloseFile(f.file); err != nil {
return fmt.Errorf("close file: %s", err)
}
}
f.file = nil
f.storage.OpWriter = nil
return nil
}
// closeStorage attempts to close storage, including unmapping the old
// storage if includeMap is true. This would normally make sense if you're
// expecting to be done using the fragment, or to reload it. But it's also
// okay to just leave stuff mmapped; you don't have to keep the file
// descriptor open. So in some cases, we'll just leave the old mmapping
// in place, rather than regenerating everything from the new file.
func (f *fragment) closeStorage(includeMap bool) error {
// Clear the storage bitmap so it doesn't access the closed mmap.
//f.storage = roaring.NewBitmap()
// Unmap the file.
if includeMap && f.storageData != nil {
if err := syswrap.Munmap(f.storageData); err != nil {
return fmt.Errorf("munmap: %s", err)
}
f.storageData = nil
}
if err := f.safeClose(); err != nil {
return err
}
// opN is determined by how many bit set/clear operations are in the storage // opN is determined by how many bit set/clear operations are in the storage
// write log, so once the storage is closed it should be 0. Opening new // write log, so once the storage is closed it should be 0. Opening new
// storage will set opN appropriately. // storage will set opN appropriately.
f.opN = 0 f.opN = 0
if f.gen != nil {
f.gen.Done()
}
return nil return nil
} }
@ -647,22 +474,17 @@ func (f *fragment) rowFromStorage(rowID uint64) *Row {
func (f *fragment) setBit(rowID, columnID uint64) (changed bool, err error) { func (f *fragment) setBit(rowID, columnID uint64) (changed bool, err error) {
f.mu.Lock() f.mu.Lock()
defer f.mu.Unlock() defer f.mu.Unlock()
mustClose, err := f.reopen() err = f.gen.Transaction(&f.storage.OpWriter, func() error {
if err != nil {
return false, errors.Wrap(err, "reopening")
}
if mustClose {
defer f.safeClose()
}
// handle mutux field type // handle mutux field type
if f.mutexVector != nil { if f.mutexVector != nil {
if err := f.handleMutex(rowID, columnID); err != nil { if err := f.handleMutex(rowID, columnID); err != nil {
return changed, errors.Wrap(err, "handling mutex") return errors.Wrap(err, "handling mutex")
} }
} }
changed, err = f.unprotectedSetBit(rowID, columnID)
return f.unprotectedSetBit(rowID, columnID) return err
})
return changed, err
} }
// handleMutex will clear an existing row and store the new row // handleMutex will clear an existing row and store the new row
@ -726,17 +548,14 @@ func (f *fragment) unprotectedSetBit(rowID, columnID uint64) (changed bool, err
// clearBit clears a bit for a given column & row within the fragment. // clearBit clears a bit for a given column & row within the fragment.
// This updates both the on-disk storage and the in-cache bitmap. // This updates both the on-disk storage and the in-cache bitmap.
func (f *fragment) clearBit(rowID, columnID uint64) (bool, error) { func (f *fragment) clearBit(rowID, columnID uint64) (changed bool, err error) {
f.mu.Lock() f.mu.Lock()
defer f.mu.Unlock() defer f.mu.Unlock()
mustClose, err := f.reopen() err = f.gen.Transaction(&f.storage.OpWriter, func() error {
if err != nil { changed, err = f.unprotectedClearBit(rowID, columnID)
return false, errors.Wrap(err, "reopening") return err
} })
if mustClose { return changed, err
defer f.safeClose()
}
return f.unprotectedClearBit(rowID, columnID)
} }
// unprotectedClearBit TODO should be replaced by an invocation of // unprotectedClearBit TODO should be replaced by an invocation of
@ -782,17 +601,14 @@ func (f *fragment) unprotectedClearBit(rowID, columnID uint64) (changed bool, er
// setRow replaces an existing row (specified by rowID) with the given // setRow replaces an existing row (specified by rowID) with the given
// Row. This updates both the on-disk storage and the in-cache bitmap. // Row. This updates both the on-disk storage and the in-cache bitmap.
func (f *fragment) setRow(row *Row, rowID uint64) (bool, error) { func (f *fragment) setRow(row *Row, rowID uint64) (changed bool, err error) {
f.mu.Lock() f.mu.Lock()
defer f.mu.Unlock() defer f.mu.Unlock()
mustClose, err := f.reopen() err = f.gen.Transaction(&f.storage.OpWriter, func() error {
if err != nil { changed, err = f.unprotectedSetRow(row, rowID)
return false, errors.Wrap(err, "reopening") return err
} })
if mustClose { return changed, err
defer f.safeClose()
}
return f.unprotectedSetRow(row, rowID)
} }
func (f *fragment) unprotectedSetRow(row *Row, rowID uint64) (changed bool, err error) { func (f *fragment) unprotectedSetRow(row *Row, rowID uint64) (changed bool, err error) {
@ -833,7 +649,7 @@ func (f *fragment) unprotectedSetRow(row *Row, rowID uint64) (changed bool, err
f.rowCache.Add(rowID, nil) f.rowCache.Add(rowID, nil)
// Snapshot storage. // Snapshot storage.
f.enqueueSnapshot() f.snapshotQueue.Enqueue(f)
f.stats.Count("setRow", 1, 1.0) f.stats.Count("setRow", 1, 1.0)
return changed, nil return changed, nil
@ -841,17 +657,14 @@ func (f *fragment) unprotectedSetRow(row *Row, rowID uint64) (changed bool, err
// ClearRow clears a row for a given rowID within the fragment. // ClearRow clears a row for a given rowID within the fragment.
// This updates both the on-disk storage and the in-cache bitmap. // This updates both the on-disk storage and the in-cache bitmap.
func (f *fragment) clearRow(rowID uint64) (bool, error) { func (f *fragment) clearRow(rowID uint64) (changed bool, err error) {
f.mu.Lock() f.mu.Lock()
defer f.mu.Unlock() defer f.mu.Unlock()
mustClose, err := f.reopen() err = f.gen.Transaction(&f.storage.OpWriter, func() error {
if err != nil { changed, err = f.unprotectedClearRow(rowID)
return false, errors.Wrap(err, "reopening") return err
} })
if mustClose { return changed, err
defer f.safeClose()
}
return f.unprotectedClearRow(rowID)
} }
func (f *fragment) unprotectedClearRow(rowID uint64) (changed bool, err error) { func (f *fragment) unprotectedClearRow(rowID uint64) (changed bool, err error) {
@ -877,7 +690,7 @@ func (f *fragment) unprotectedClearRow(rowID uint64) (changed bool, err error) {
f.rowCache.Add(rowID, nil) f.rowCache.Add(rowID, nil)
// Snapshot storage. // Snapshot storage.
f.enqueueSnapshot() f.snapshotQueue.Enqueue(f)
f.stats.Count("clearRow", 1, 1.0) f.stats.Count("clearRow", 1, 1.0)
@ -977,14 +790,7 @@ func (f *fragment) positionsForValue(columnID uint64, bitDepth uint, value int64
func (f *fragment) setValueBase(columnID uint64, bitDepth uint, value int64, clear bool) (changed bool, err error) { func (f *fragment) setValueBase(columnID uint64, bitDepth uint, value int64, clear bool) (changed bool, err error) {
f.mu.Lock() f.mu.Lock()
defer f.mu.Unlock() defer f.mu.Unlock()
mustClose, err := f.reopen() err = f.gen.Transaction(&f.storage.OpWriter, func() error {
if err != nil {
return false, errors.Wrap(err, "reopening")
}
if mustClose {
defer f.safeClose()
}
// Convert value to an unsigned representation. // Convert value to an unsigned representation.
uvalue := uint64(value) uvalue := uint64(value)
if value < 0 { if value < 0 {
@ -994,13 +800,13 @@ func (f *fragment) setValueBase(columnID uint64, bitDepth uint, value int64, cle
for i := uint(0); i < bitDepth; i++ { for i := uint(0); i < bitDepth; i++ {
if uvalue&(1<<i) != 0 { if uvalue&(1<<i) != 0 {
if c, err := f.unprotectedSetBit(uint64(bsiOffsetBit+i), columnID); err != nil { if c, err := f.unprotectedSetBit(uint64(bsiOffsetBit+i), columnID); err != nil {
return changed, err return err
} else if c { } else if c {
changed = true changed = true
} }
} else { } else {
if c, err := f.unprotectedClearBit(uint64(bsiOffsetBit+i), columnID); err != nil { if c, err := f.unprotectedClearBit(uint64(bsiOffsetBit+i), columnID); err != nil {
return changed, err return err
} else if c { } else if c {
changed = true changed = true
} }
@ -1010,13 +816,13 @@ func (f *fragment) setValueBase(columnID uint64, bitDepth uint, value int64, cle
// Mark value as set (or cleared). // Mark value as set (or cleared).
if clear { if clear {
if c, err := f.unprotectedClearBit(uint64(bsiExistsBit), columnID); err != nil { if c, err := f.unprotectedClearBit(uint64(bsiExistsBit), columnID); err != nil {
return changed, errors.Wrap(err, "clearing not-null") return errors.Wrap(err, "clearing not-null")
} else if c { } else if c {
changed = true changed = true
} }
} else { } else {
if c, err := f.unprotectedSetBit(uint64(bsiExistsBit), columnID); err != nil { if c, err := f.unprotectedSetBit(uint64(bsiExistsBit), columnID); err != nil {
return changed, errors.Wrap(err, "marking not-null") return errors.Wrap(err, "marking not-null")
} else if c { } else if c {
changed = true changed = true
} }
@ -1025,19 +831,21 @@ func (f *fragment) setValueBase(columnID uint64, bitDepth uint, value int64, cle
// Mark sign bit (or clear). // Mark sign bit (or clear).
if value >= 0 || clear { if value >= 0 || clear {
if c, err := f.unprotectedClearBit(uint64(bsiSignBit), columnID); err != nil { if c, err := f.unprotectedClearBit(uint64(bsiSignBit), columnID); err != nil {
return changed, errors.Wrap(err, "clearing sign") return errors.Wrap(err, "clearing sign")
} else if c { } else if c {
changed = true changed = true
} }
} else { } else {
if c, err := f.unprotectedSetBit(uint64(bsiSignBit), columnID); err != nil { if c, err := f.unprotectedSetBit(uint64(bsiSignBit), columnID); err != nil {
return changed, errors.Wrap(err, "marking sign") return errors.Wrap(err, "marking sign")
} else if c { } else if c {
changed = true changed = true
} }
} }
return changed, nil return nil
})
return changed, err
} }
// importSetValue is a more efficient SetValue just for imports. // importSetValue is a more efficient SetValue just for imports.
@ -1353,10 +1161,25 @@ func (f *fragment) rangeLT(bitDepth uint, predicate int64, allowEquality bool) (
return f.rangeGTUnsigned(b.Intersect(f.row(bsiSignBit)), bitDepth, upredicate, allowEquality) return f.rangeGTUnsigned(b.Intersect(f.row(bsiSignBit)), bitDepth, upredicate, allowEquality)
} }
// msb gives the 1-indexed position (counting from lsb) of the most
// significant bit. E.G. for 1 it would return 1, for 2 2, for 3 2,
// for 4 3, for 8 4, etc.
func msb(x uint64) uint {
lz := bits.LeadingZeros64(x)
return 64 - uint(lz)
}
// rangeLTUnsigned returns all bits LT/LTE the predicate without considering the sign bit. // rangeLTUnsigned returns all bits LT/LTE the predicate without considering the sign bit.
func (f *fragment) rangeLTUnsigned(filter *Row, bitDepth uint, predicate uint64, allowEquality bool) (*Row, error) { func (f *fragment) rangeLTUnsigned(filter *Row, bitDepth uint, predicate uint64, allowEquality bool) (*Row, error) {
keep := NewRow() keep := NewRow()
// if the predicate is larger than all representable numbers given
// our bitDepth... then just return everything.
if msb(predicate) > bitDepth {
return filter, nil
}
// Filter any bits that don't match the current bit value. // Filter any bits that don't match the current bit value.
leadingZeros := true leadingZeros := true
for i := int(bitDepth - 1); i >= 0; i-- { for i := int(bitDepth - 1); i >= 0; i-- {
@ -2051,14 +1874,7 @@ func (f *fragment) bulkImportStandard(rowIDs, columnIDs []uint64, options *Impor
// snapshot of the fragment or just do in-memory updates while appending // snapshot of the fragment or just do in-memory updates while appending
// operations to the op log. // operations to the op log.
func (f *fragment) importPositions(set, clear []uint64, rowSet map[uint64]struct{}) error { func (f *fragment) importPositions(set, clear []uint64, rowSet map[uint64]struct{}) error {
mustClose, err := f.reopen() err := f.gen.Transaction(&f.storage.OpWriter, func() error {
if err != nil {
return errors.Wrap(err, "reopening")
}
if mustClose {
defer f.safeClose()
}
if len(set) > 0 { if len(set) > 0 {
f.stats.Count("ImportingN", int64(len(set)), 1) f.stats.Count("ImportingN", int64(len(set)), 1)
changedN, err := f.storage.AddN(set...) // TODO benchmark Add/RemoveN behavior with sorted/unsorted positions changedN, err := f.storage.AddN(set...) // TODO benchmark Add/RemoveN behavior with sorted/unsorted positions
@ -2095,8 +1911,9 @@ func (f *fragment) importPositions(set, clear []uint64, rowSet map[uint64]struct
if f.CacheType != CacheTypeNone { if f.CacheType != CacheTypeNone {
f.cache.Recalculate() f.cache.Recalculate()
} }
return nil return nil
})
return err
} }
// bulkImportMutex performs a bulk import on a fragment while ensuring // bulkImportMutex performs a bulk import on a fragment while ensuring
@ -2182,7 +1999,6 @@ func (f *fragment) importValueSmallWrite(columnIDs []uint64, values []int64, bit
} }
return nil return nil
}(); err != nil { }(); err != nil {
_ = f.closeStorage(true)
_ = f.openStorage(true) _ = f.openStorage(true)
return err return err
} }
@ -2191,7 +2007,14 @@ func (f *fragment) importValueSmallWrite(columnIDs []uint64, values []int64, bit
rowSet[uint64(i)] = struct{}{} rowSet[uint64(i)] = struct{}{}
} }
err := f.importPositions(toSet, toClear, rowSet) err := f.importPositions(toSet, toClear, rowSet)
if err != nil {
return errors.Wrap(err, "importing positions") return errors.Wrap(err, "importing positions")
}
// Reset the rowCache.
f.rowCache = &simpleCache{make(map[uint64]*Row)}
return nil
} }
// importValue bulk imports a set of range-encoded values. // importValue bulk imports a set of range-encoded values.
@ -2223,20 +2046,23 @@ func (f *fragment) importValue(columnIDs []uint64, values []int64, bitDepth uint
} }
return nil return nil
}(); err != nil { }(); err != nil {
_ = f.closeStorage(true)
_ = f.openStorage(true) _ = f.openStorage(true)
return err return err
} }
// We don't actually care, except we want our stats to be accurate. // Keep stats accurate. We don't call incrementOpN here because it may
f.incrementOpN(totalChanges) // or may not enqueue a request, which would then be in the queue
// taking up space and otherwise being a possible nuisance, when we're
// about to force a snapshot anyway.
f.opN += totalChanges
f.ops++
// in theory, this should probably have happened anyway, but if enough // Reset the rowCache.
f.rowCache = &simpleCache{make(map[uint64]*Row)}
// in theory, this should probably have been queued anyway, but if enough
// of the bits matched existing bits, we'll be under our opN estimate, and // of the bits matched existing bits, we'll be under our opN estimate, and
// we want to ensure that the snapshot happens. // we want to ensure that the snapshot happens.
f.enqueueSnapshot() return f.snapshotQueue.Immediate(f)
f.unprotectedAwaitSnapshot()
return nil
} }
// importRoaring imports from the official roaring data format defined at // importRoaring imports from the official roaring data format defined at
@ -2251,7 +2077,13 @@ func (f *fragment) importRoaring(ctx context.Context, data []byte, clear bool) e
defer f.mu.Unlock() defer f.mu.Unlock()
span.Finish() span.Finish()
span, ctx = tracing.StartSpanFromContext(ctx, "importRoaring.ImportRoaringBits") span, ctx = tracing.StartSpanFromContext(ctx, "importRoaring.ImportRoaringBits")
changed, rowSet, err := f.storage.ImportRoaringBits(data, clear, true, rowSize) var changed int
var rowSet map[uint64]int
err := f.gen.Transaction(&f.storage.OpWriter, func() (err error) {
changed, rowSet, err = f.storage.ImportRoaringBits(data, clear, true, rowSize)
return err
})
span.Finish() span.Finish()
if err != nil { if err != nil {
return err return err
@ -2267,9 +2099,18 @@ func (f *fragment) importRoaring(ctx context.Context, data []byte, clear bool) e
f.rowCache.Add(rowID, nil) f.rowCache.Add(rowID, nil)
if updateCache { if updateCache {
anyChanged = true anyChanged = true
if changes < 0 {
absChanges := uint64(-1 * changes)
if absChanges <= f.cache.Get(rowID) {
f.cache.BulkAdd(rowID, f.cache.Get(rowID)-absChanges)
} else {
f.cache.BulkAdd(rowID, 0)
}
} else {
f.cache.BulkAdd(rowID, f.cache.Get(rowID)+uint64(changes)) f.cache.BulkAdd(rowID, f.cache.Get(rowID)+uint64(changes))
} }
} }
}
// we only set this if we need to update the cache // we only set this if we need to update the cache
if anyChanged { if anyChanged {
f.cache.Recalculate() f.cache.Recalculate()
@ -2290,14 +2131,14 @@ func (f *fragment) incrementOpN(changed int) {
f.opN += changed f.opN += changed
f.ops++ f.ops++
if f.opN > f.MaxOpN { if f.opN > f.MaxOpN {
f.enqueueSnapshot() f.snapshotQueue.Enqueue(f)
} }
} }
// Snapshot writes the storage bitmap to disk and reopens it. This may // Snapshot writes the storage bitmap to disk and reopens it. This may
// coexist with existing background-queue snapshotting; it does not remove // coexist with existing background-queue snapshotting; it does not remove
// things from the queue. You probably don't want to do this; use // things from the queue. You probably don't want to do this; use
// enqueueSnapshot/awaitSnapshot. // the snapshotQueue's Enqueue/Await.
func (f *fragment) Snapshot() error { func (f *fragment) Snapshot() error {
f.mu.Lock() f.mu.Lock()
defer f.mu.Unlock() defer f.mu.Unlock()
@ -2310,25 +2151,13 @@ func track(start time.Time, message string, stats stats.StatsClient, logger logg
stats.Histogram("snapshot", elapsed.Seconds(), 1.0) stats.Histogram("snapshot", elapsed.Seconds(), 1.0)
} }
// protectedSnapshot grabs the lock and unconditionally calls snapshot(). If
// fromQueue is true, the snapshotting state is also cleared.
func (f *fragment) protectedSnapshot(fromQueue bool) error {
f.mu.Lock()
defer f.mu.Unlock()
err := f.snapshot()
if fromQueue {
f.snapshotting = false
}
return err
}
// snapshot does the actual snapshot operation. it does not check or care // snapshot does the actual snapshot operation. it does not check or care
// about f.snapshotting. // about f.snapshotPending.
func (f *fragment) snapshot() error { func (f *fragment) snapshot() error {
f.totalOpN += int64(f.opN)
f.totalOps += int64(f.ops)
f.snapshotsTaken++
_, err := unprotectedWriteToFragment(f, f.storage) _, err := unprotectedWriteToFragment(f, f.storage)
if err == nil {
f.snapshotStamp = time.Now()
}
return err return err
} }
@ -2345,22 +2174,24 @@ func unprotectedWriteToFragment(f *fragment, bm *roaring.Bitmap) (n int64, err e
if err != nil { if err != nil {
return n, fmt.Errorf("create snapshot file: %s", err) return n, fmt.Errorf("create snapshot file: %s", err)
} }
defer file.Close() // No deferred close, because we want to close it sooner than the
// end of this function.
// Write storage to snapshot. // Write storage to snapshot.
bw := bufio.NewWriter(file) bw := bufio.NewWriter(file)
if n, err = bm.WriteTo(bw); err != nil { if n, err = bm.WriteTo(bw); err != nil {
file.Close()
return n, fmt.Errorf("snapshot write to: %s", err) return n, fmt.Errorf("snapshot write to: %s", err)
} }
if err := bw.Flush(); err != nil { if err := bw.Flush(); err != nil {
file.Close()
return n, fmt.Errorf("flush: %s", err) return n, fmt.Errorf("flush: %s", err)
} }
// Close current storage. // we close the file here so we don't still have it open when trying
if err := f.closeStorage(false); err != nil { // to open it in a moment.
return n, fmt.Errorf("close storage: %s", err) file.Close()
}
// Move snapshot to data file location. // Move snapshot to data file location.
if err := os.Rename(snapshotPath, f.path); err != nil { if err := os.Rename(snapshotPath, f.path); err != nil {
@ -2560,11 +2391,6 @@ func (f *fragment) readStorageFromArchive(r io.Reader) error {
return errors.Wrap(err, "copying") return errors.Wrap(err, "copying")
} }
// Close current storage.
if err := f.closeStorage(true); err != nil {
return errors.Wrap(err, "closing")
}
// Move snapshot to data file location. // Move snapshot to data file location.
if err := os.Rename(path, f.path); err != nil { if err := os.Rename(path, f.path); err != nil {
return errors.Wrap(err, "renaming") return errors.Wrap(err, "renaming")

View file

@ -596,6 +596,21 @@ func TestFragment_Range(t *testing.T) {
} }
}) })
t.Run("LTRegression", func(t *testing.T) {
f := mustOpenFragment("i", "f", viewStandard, 0, "")
defer f.Clean(t)
if _, err := f.setValue(1, 1, 1); err != nil {
t.Fatal(err)
}
if b, err := f.rangeOp(pql.LT, 1, 2); err != nil {
t.Fatal(err)
} else if !reflect.DeepEqual(b.Columns(), []uint64{1}) {
t.Fatalf("unepxected coulmns: %+v", b.Columns())
}
})
t.Run("GT", func(t *testing.T) { t.Run("GT", func(t *testing.T) {
f := mustOpenFragment("i", "f", viewStandard, 0, "") f := mustOpenFragment("i", "f", viewStandard, 0, "")
defer f.Clean(t) defer f.Clean(t)
@ -1382,6 +1397,7 @@ func TestFragment_WriteTo_ReadFrom(t *testing.T) {
// Read into another fragment. // Read into another fragment.
f1 := mustOpenFragment("i", "f", viewStandard, 0, "") f1 := mustOpenFragment("i", "f", viewStandard, 0, "")
defer f1.Clean(t)
if rn, err := f1.ReadFrom(&buf); err != nil { if rn, err := f1.ReadFrom(&buf); err != nil {
t.Fatal(err) t.Fatal(err)
} else if wn != rn { } else if wn != rn {
@ -2054,7 +2070,11 @@ func BenchmarkImportRoaring(b *testing.B) {
b.StartTimer() b.StartTimer()
err := f.importRoaringT(data, false) err := f.importRoaringT(data, false)
if err != nil { if err != nil {
f.awaitSnapshot() // we don't actually particularly
// care whether this succeeds,
// but if it's happening we want
// it to be done.
_ = f.snapshotQueue.Await(f)
f.Clean(b) f.Clean(b)
b.Fatalf("import error: %v", err) b.Fatalf("import error: %v", err)
} }
@ -2093,7 +2113,9 @@ func BenchmarkImportRoaringConcurrent(b *testing.B) {
j := j j := j
eg.Go(func() error { eg.Go(func() error {
err := frags[j].importRoaringT(data[j], false) err := frags[j].importRoaringT(data[j], false)
frags[j].awaitSnapshot() // error unimportant if it happened, but we want
// any snapshots to have finished.
_ = frags[j].snapshotQueue.Await(frags[j])
return err return err
}) })
} }
@ -2131,11 +2153,13 @@ func BenchmarkImportRoaringUpdateConcurrent(b *testing.B) {
// is excessive. force storage into snapshotted state, then use import // is excessive. force storage into snapshotted state, then use import
// to generate an op log and/or snapshot. // to generate an op log and/or snapshot.
_, _, err := frags[j].storage.ImportRoaringBits(data, false, false, 0) _, _, err := frags[j].storage.ImportRoaringBits(data, false, false, 0)
frags[j].enqueueSnapshot()
frags[j].awaitSnapshot()
if err != nil { if err != nil {
b.Fatalf("importing roaring: %v", err) b.Fatalf("importing roaring: %v", err)
} }
err = frags[j].snapshotQueue.Immediate(frags[j])
if err != nil {
b.Fatalf("snapshot after import: %v", err)
}
} }
eg := errgroup.Group{} eg := errgroup.Group{}
b.StartTimer() b.StartTimer()
@ -2143,7 +2167,10 @@ func BenchmarkImportRoaringUpdateConcurrent(b *testing.B) {
j := j j := j
eg.Go(func() error { eg.Go(func() error {
err := frags[j].importRoaringT(updata, false) err := frags[j].importRoaringT(updata, false)
frags[j].awaitSnapshot() err2 := frags[j].snapshotQueue.Await(frags[j])
if err == nil {
err = err2
}
return err return err
}) })
} }
@ -2205,20 +2232,38 @@ func BenchmarkImportRoaringUpdate(b *testing.B) {
// is excessive. force storage into snapshotted state, then use import // is excessive. force storage into snapshotted state, then use import
// to generate an op log and/or snapshot. // to generate an op log and/or snapshot.
_, _, err := f.storage.ImportRoaringBits(data, false, false, 0) _, _, err := f.storage.ImportRoaringBits(data, false, false, 0)
f.enqueueSnapshot()
f.awaitSnapshot()
if err != nil { if err != nil {
b.Errorf("import error: %v", err) b.Errorf("import error: %v", err)
} }
err = f.snapshotQueue.Immediate(f)
if err != nil {
b.Errorf("snapshot after import error: %v", err)
}
b.StartTimer() b.StartTimer()
err = f.importRoaringT(updata, false) err = f.importRoaringT(updata, false)
f.awaitSnapshot()
if err != nil { if err != nil {
f.Clean(b) f.Clean(b)
b.Errorf("import error: %v", err) b.Errorf("import error: %v", err)
} }
err = f.snapshotQueue.Await(f)
if err != nil {
b.Errorf("snapshot after import error: %v", err)
}
b.StopTimer() b.StopTimer()
stat, _ := f.file.Stat() var stat os.FileInfo
var statTarget io.Writer
err = f.gen.Transaction(&statTarget, func() error {
targetFile, ok := statTarget.(*os.File)
if ok {
stat, _ = targetFile.Stat()
} else {
b.Errorf("couldn't stat file")
}
return nil
})
if err != nil {
b.Errorf("transaction error: %v", err)
}
fileSize[name] = stat.Size() fileSize[name] = stat.Size()
f.Clean(b) f.Clean(b)
} }
@ -2367,6 +2412,7 @@ func TestGetZipfRowsSliceRoaring(t *testing.T) {
t.Fatalf("suspect distribution from getZipfRowsSliceRoaring") t.Fatalf("suspect distribution from getZipfRowsSliceRoaring")
} }
} }
f.Clean(t)
} }
// getZipfRowsSliceRoaring generates a random fragment with the given number of // getZipfRowsSliceRoaring generates a random fragment with the given number of
@ -2512,16 +2558,28 @@ func (f *fragment) sanityCheck(t testing.TB) {
} }
func (f *fragment) Clean(t testing.TB) { func (f *fragment) Clean(t testing.TB) {
f.awaitSnapshot() f.mu.Lock()
err := f.snapshotQueue.Await(f)
f.mu.Unlock()
if err != nil {
t.Fatalf("snapshot failed before sanity check: %v", err)
}
f.sanityCheck(t) f.sanityCheck(t)
if f.storage != nil && f.storage.Source != nil {
if f.storage.Source.Dead() {
t.Fatalf("cleaning up fragment %s, source %s, source already dead", f.path, f.storage.Source.ID())
}
}
errc := f.Close() errc := f.Close()
// prevent double-closes of generation during testing.
f.gen = nil
errf := os.Remove(f.path) errf := os.Remove(f.path)
errp := os.Remove(f.cachePath()) errp := os.Remove(f.cachePath())
if errc != nil || errf != nil { if errc != nil || errf != nil {
t.Fatal("cleaning up fragment: ", errc, errf, errp) t.Fatal("cleaning up fragment: ", errc, errf, errp)
} }
if f.snapshotQueue != nil { if f.snapshotQueue != nil {
close(f.snapshotQueue) f.snapshotQueue.Stop()
f.snapshotQueue = nil f.snapshotQueue = nil
} }
// not all fragments have cache files // not all fragments have cache files
@ -2544,7 +2602,7 @@ func (f *fragment) CleanKeep(t testing.TB) {
t.Fatal("closing fragment: ", errc, errp) t.Fatal("closing fragment: ", errc, errp)
} }
if f.snapshotQueue != nil { if f.snapshotQueue != nil {
close(f.snapshotQueue) f.snapshotQueue.Stop()
f.snapshotQueue = nil f.snapshotQueue = nil
} }
// not all fragments have cache files // not all fragments have cache files
@ -3032,10 +3090,10 @@ func TestFragmentRowIterator(t *testing.T) {
func TestUnionInPlaceMapped(t *testing.T) { func TestUnionInPlaceMapped(t *testing.T) {
f := mustOpenFragment("i", "f", "v", 0, CacheTypeNone) f := mustOpenFragment("i", "f", "v", 0, CacheTypeNone)
// note: clean has to be deferred first, because it has to run with
// the lock *not* held, because it is sometimes so it has to grab the
// lock...
defer f.Clean(t) defer f.Clean(t)
// I know this doesn't actually matter in our current context, but
// strictly speaking, we do say you have to hold the lock while calling
// unprotectedWriteToFragment...
f.mu.Lock() f.mu.Lock()
defer f.mu.Unlock() defer f.mu.Unlock()
r0 := rand.New(rand.NewSource(2)) r0 := rand.New(rand.NewSource(2))
@ -3068,8 +3126,15 @@ func TestUnionInPlaceMapped(t *testing.T) {
f.storage.UnionInPlace(setBM1) f.storage.UnionInPlace(setBM1)
countUnion := f.storage.Count() countUnion := f.storage.Count()
// UnionInPlace produces no ops log, we have to make it snapshot, to // UnionInPlace produces no ops log, we have to make it snapshot, to
// ensure that the on-disk representation is correct. // ensure that the on-disk representation is correct. Note, UIP is
f.enqueueSnapshot() // not used for things that are modifying real fragments, usually;
// it's used only in computation of things that usually don't go to
// disk, which is why we handle this specially in testing and not
// generically.
err = f.snapshotQueue.Immediate(f)
if err != nil {
t.Fatalf("snapshot after union-in-place: %v", err)
}
if count0 != countF { if count0 != countF {
t.Fatalf("writing bitmap to storage changed count: %d => %d", count0, countF) t.Fatalf("writing bitmap to storage changed count: %d => %d", count0, countF)
@ -3178,6 +3243,25 @@ func TestFragmentPositionsForValue(t *testing.T) {
} }
} }
func TestIntLTRegression(t *testing.T) {
f := mustOpenFragment("i", "f", "v", 0, CacheTypeNone)
defer f.Clean(t)
_, err := f.setValue(1, 6, 33)
if err != nil {
t.Fatalf("setting value: %v", err)
}
row, err := f.rangeOp(pql.LT, 6, 33)
if err != nil {
t.Fatalf("doing range of: %v", err)
}
if !row.IsEmpty() {
t.Errorf("expected nothing, but got: %v", row.Columns())
}
}
func TestImportClearRestart(t *testing.T) { func TestImportClearRestart(t *testing.T) {
tests := []struct { tests := []struct {
rows []uint64 rows []uint64
@ -3257,7 +3341,7 @@ func TestImportClearRestart(t *testing.T) {
f2.MaxOpN = maxOpN f2.MaxOpN = maxOpN
f2.CacheType = f.CacheType f2.CacheType = f.CacheType
err = f.closeStorage(true) err = f.closeStorage()
if err != nil { if err != nil {
t.Fatalf("closing storage: %v", err) t.Fatalf("closing storage: %v", err)
} }
@ -3291,7 +3375,7 @@ func TestImportClearRestart(t *testing.T) {
f3.MaxOpN = maxOpN f3.MaxOpN = maxOpN
f3.CacheType = f.CacheType f3.CacheType = f.CacheType
err = f2.closeStorage(true) err = f2.closeStorage()
if err != nil { if err != nil {
t.Fatalf("f2 closing storage: %v", err) t.Fatalf("f2 closing storage: %v", err)
} }
@ -3339,6 +3423,7 @@ func check(t *testing.T, f *fragment, exp map[uint64]map[uint64]struct{}) {
func TestImportValueConcurrent(t *testing.T) { func TestImportValueConcurrent(t *testing.T) {
f := mustOpenBSIFragment("i", "f", viewBSIGroupPrefix+"foo", 0) f := mustOpenBSIFragment("i", "f", viewBSIGroupPrefix+"foo", 0)
defer f.Clean(t)
eg := &errgroup.Group{} eg := &errgroup.Group{}
for i := 0; i < 4; i++ { for i := 0; i < 4; i++ {
i := i i := i
@ -3363,7 +3448,7 @@ func TestImportMultipleValues(t *testing.T) {
cols []uint64 cols []uint64
vals []int64 vals []int64
checkCols []uint64 checkCols []uint64
checkVals []uint64 checkVals []int64
depth uint depth uint
}{ }{
{ {
@ -3371,7 +3456,7 @@ func TestImportMultipleValues(t *testing.T) {
vals: []int64{97, 100}, vals: []int64{97, 100},
depth: 7, depth: 7,
checkCols: []uint64{0}, checkCols: []uint64{0},
checkVals: []uint64{100}, checkVals: []int64{100},
}, },
} }
@ -3395,7 +3480,7 @@ func TestImportMultipleValues(t *testing.T) {
if !exists { if !exists {
t.Errorf("column %d should exist", cc) t.Errorf("column %d should exist", cc)
} }
if n != 100 { if n != cv {
t.Errorf("wrong value: %d is not %d", n, cv) t.Errorf("wrong value: %d is not %d", n, cv)
} }
} }
@ -3405,6 +3490,66 @@ func TestImportMultipleValues(t *testing.T) {
} }
} }
func TestImportValueRowCache(t *testing.T) {
type testCase struct {
cols []uint64
vals []int64
checkCols []uint64
depth uint
}
tests := []struct {
tc1 testCase
tc2 testCase
}{
{
tc1: testCase{
cols: []uint64{2},
vals: []int64{1},
depth: 1,
checkCols: []uint64{2},
},
tc2: testCase{
cols: []uint64{1000},
vals: []int64{1},
depth: 1,
checkCols: []uint64{2, 1000},
},
},
}
for i, test := range tests {
for _, maxOpN := range []int{1, 10000} {
t.Run(fmt.Sprintf("%dMaxOpN%d", i, maxOpN), func(t *testing.T) {
f := mustOpenBSIFragment("i", "f", viewBSIGroupPrefix+"foo", 0)
f.MaxOpN = maxOpN
defer f.Clean(t)
// First import (tc1)
if err := f.importValue(test.tc1.cols, test.tc1.vals, test.tc1.depth, false); err != nil {
t.Fatalf("importing values: %v", err)
}
if r, err := f.rangeOp(pql.GT, test.tc1.depth, 0); err != nil {
t.Error("getting range of values")
} else if !reflect.DeepEqual(r.Columns(), test.tc1.checkCols) {
t.Errorf("wrong column values. expected: %v, but got: %v", test.tc1.checkCols, r.Columns())
}
// Second import (tc2)
if err := f.importValue(test.tc2.cols, test.tc2.vals, test.tc2.depth, false); err != nil {
t.Fatalf("importing values: %v", err)
}
if r, err := f.rangeOp(pql.GT, test.tc2.depth, 0); err != nil {
t.Error("getting range of values")
} else if !reflect.DeepEqual(r.Columns(), test.tc2.checkCols) {
t.Errorf("wrong column values. expected: %v, but got: %v", test.tc2.checkCols, r.Columns())
}
})
}
}
}
func TestFragmentConcurrentReadWrite(t *testing.T) { func TestFragmentConcurrentReadWrite(t *testing.T) {
f := mustOpenFragment("i", "f", viewStandard, 0, CacheTypeRanked) f := mustOpenFragment("i", "f", viewStandard, 0, CacheTypeRanked)
defer f.Clean(t) defer f.Clean(t)

407
generation.go Normal file
View file

@ -0,0 +1,407 @@
// Copyright 2019 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package pilosa
import (
"fmt"
"io"
"io/ioutil"
"os"
"runtime"
"sync"
"syscall"
"time"
"github.com/pilosa/pilosa/v2/logger"
"github.com/pilosa/pilosa/v2/roaring"
"github.com/pilosa/pilosa/v2/syswrap"
"github.com/pkg/errors"
)
// generation represents one "generation" of opening a data file.
// This is what determines when it's safe to unmap a data file, if it
// got mapped, and handles closing/reopening files if we need to
// manage file handle availability. It's an interface because this
// lets us write simpler code for specific cases, rather than handling
// the whole matrix of mapped/unmapped, staying open/being reopened,
// etcetera.
//
// You create a generation by calling newGeneration with a file
// path. If it succeeds in opening that path, it calls a provided
// setup function with the data from the generation, and a flag
// indicating whether the data is mmapped. If the setup function
// fails, newGeneration cleans things up and closes. Otherwise,
// it returns a generation.
//
// The generation itself uses runtime.SetFinalizer to clean up when
// the last reference to it goes away. You should store a pointer
// to the generation in any object which is reliant on the generation.
//
// When you anticipate a generation should be done (for instance,
// opening a new generation), the old one gets marked done, which
// stashes a timestamp in it. Later operations can check whether
// the timestamp is a while back, and if so, complain that something
// might be wrong.
//
// In some cases, we don't have enough open file limit to keep every
// file actually open. To address this, use the `Transaction` function,
// which ensures that the file is open, stores a reference to it in
// a provided `*io.Writer`, and then restores the previous value of
// the io.Writer when it's done. For instance, for a bitmap, this might
// be used with `&b.OpWriter`.
//
// newGeneration takes an optional previous generation; it calls
// that generation's Done function after running the provided setup,
// and bumps the generation count.
type generation interface {
// Transaction runs the given transaction with the generation's
// file open. If the **os.File parameter is
// non-nil, the generation's file will be open, and stored
// into that pointer, during the execution of func, after
// which the previous contents are restored. Otherwise
// the file may or may not be open during the operation.
Transaction(*io.Writer, func() error) error
// Done() should be called exactly once, to indicate that a
// generation is expected not to be in use for long -- for instance,
// when a new generation replaces it.
Done()
// Generation count.
Generation() int64
// ID indicates the source -- path and generation number -- that
// this generation represents.
ID() string
Dead() bool
}
type mmapGeneration struct {
mu sync.Mutex // mutex guards modifiers of generation, not of data
transMu sync.Mutex // guards transactions, specifically
path string
id string
file *os.File
data []byte
generation int64 // generation counter
dead bool // we think this generation is dead
deadSince time.Time // when this generation was marked dead
retries int // for cases where we're retrying
logger logger.Logger
}
func (m *mmapGeneration) Dead() bool {
m.mu.Lock()
defer m.mu.Unlock()
return m.dead
}
func (m *mmapGeneration) ID() string {
return m.id
}
func (m *mmapGeneration) Generation() int64 {
return m.generation
}
// Transaction runs an exclusive call, ensuring that the file is open if
// the *io.Writer parameter is present.
func (m *mmapGeneration) Transaction(fileP *io.Writer, fn func() error) (transactionErr error) {
m.transMu.Lock()
defer m.transMu.Unlock()
// HEY LOOK CAREFULLY AT THIS BIT:
// We can't just defer this unlock. We specifically want to be
// sure to unlock the regular mutex *before* this function is over,
// and if we error out trying to open the file, we want to do it
// even sooner. If we deferred this, the transaction would block
// *everything*, including things like sanity checks against the
// generation being Dead(), but also including the deferred
// re-close-the-file.
m.mu.Lock()
// if we've been asked for a file pointer, we need to ensure that
// our file is open, and that the file pointer to it is stored in
// the requested location, then revert that when we're done.
// if we aren't asked for a file pointer, nothing needs the file
// open.
if m.dead {
elapsed := time.Since(m.deadSince)
m.logger.Printf("WARNING: transaction against %s, which has been dead for %v\n", m.id, elapsed)
}
if fileP != nil {
if m.file == nil {
// we ignore the shouldClose response here; if this
// fragment was previously not being kept open, we're
// going to stick with that.
_, err := m.openFile()
if err != nil {
m.mu.Unlock()
return err
}
defer func() {
// report a close error if we have no other error to report
m.mu.Lock()
defer m.mu.Unlock()
err := m.closeFile()
if transactionErr == nil {
transactionErr = err
}
}()
}
var fileStash io.Writer
fileStash, *fileP = *fileP, m.file
defer func() {
*fileP = fileStash
}()
}
// We are done locking the generation itself for now.
m.mu.Unlock()
return fn()
}
// Done marks the generation done, and closes its file, but may not unmap it.
// It's still conceptually possible to end up doing a Transaction against a
// done generation, but it's a red flag.
func (m *mmapGeneration) Done() {
if m == nil {
return
}
m.mu.Lock()
defer m.mu.Unlock()
if m.dead {
oops := fmt.Sprintf("generation %s, marked done again at %v, previously marked dead at %v",
m.id, time.Now(), m.deadSince)
panic(oops)
}
m.dead = true
m.deadSince = time.Now()
err := m.closeFile()
if err != nil {
m.logger.Printf("error closing generation %s: %v", m.id, err)
}
// If we're not debugging, the finalizer won't have been enabled
// previously. Finalizers have non-zero cost, so having them not be
// created until they're needed seems rewarding?
if !generationDebug {
runtime.SetFinalizer(m, generationFinalizer)
}
endGeneration(m.id)
// note, Done() doesn't close the file; only the finalizer actually
// does the shutdown.
}
// Try to close the file if it's currently open.
func (m *mmapGeneration) closeFile() error {
var lastErr error
// report the most serious error encountered, but still close
// file even if something else failed.
if m.file != nil {
if err := m.file.Sync(); err != nil {
lastErr = fmt.Errorf("sync: %s", err)
}
if err := syscall.Flock(int(m.file.Fd()), syscall.LOCK_UN); err != nil {
lastErr = fmt.Errorf("unlock: %s", err)
}
if err := syswrap.CloseFile(m.file); err != nil {
lastErr = fmt.Errorf("close file: %s", err)
}
m.file = nil
}
return lastErr
}
// openFile ensures the file is open and locked, or fails. If it does
// open the file, it will also report the "you need to close this file
// when you're done" flag from syswrap.
func (m *mmapGeneration) openFile() (shouldClose bool, err error) {
if m.file != nil {
return false, nil
}
m.file, shouldClose, err = syswrap.OpenFile(m.path, os.O_RDWR|os.O_CREATE|os.O_APPEND, 0666)
if err != nil {
return false, err
}
// do we actually want this in every openFile? I don't know.
if err := syscall.Flock(int(m.file.Fd()), syscall.LOCK_EX|syscall.LOCK_NB); err != nil {
m.file.Close()
m.file = nil
return false, fmt.Errorf("flock: %s", err)
}
return shouldClose, nil
}
func generationFinalizer(m *mmapGeneration) {
m.mu.Lock()
if !m.dead {
m.logger.Printf("finalizing generation %s which isn't dead yet\n",
m.id)
}
m.mu.Unlock()
err := m.closeFile()
if err != nil {
m.logger.Printf("finalizing generation, closing file: %v\n", err)
}
if m.data != nil {
err := syswrap.Munmap(m.data)
if err != nil {
m.logger.Printf("finalizing generation, munmap: %v\n", err)
}
m.data = nil
}
finalizeGeneration(m.id)
}
// Cancel closes a generation out entirely. It cancels any finalizer,
// unmaps any data, ends generation tracking, and closes any files.
// It does each of these separately whether or not the others need to be done,
// or succeed. It's used to handle failures from newGeneration; it makes sure
// the generation isn't holding any resources and doesn't need to be cleaned
// up otherwise.
//
// Mostly a helper function because there's several cases where newGeneration
// might fail.
func (m *mmapGeneration) Cancel() {
if m.data != nil {
_ = syswrap.Munmap(m.data)
m.data = nil
}
err := m.closeFile()
if err != nil {
m.logger.Printf("error cancelling generation %s: %v", m.id, err)
}
runtime.SetFinalizer(m, nil)
m.dead = true
m.deadSince = time.Now()
cancelGeneration(m.id)
}
// newGeneration creates a new generation using the given file path. It
// then calls the provided setup function with the allocated storage, a
// file handle, the new generation, and a flag indicatting whether the storage
// is memory-mapped. If the setup function returns a non-nil error, the
// generation is cleaned up, and newGeneration fails. The setup function
// also returns a boolean indicating whether it used the mapping; if it
// didn't, newGeneration discards the mapping and returns a nil generation.
//
// If generationDebug is enabled, we track the generation even if no mapping
// is actually in use, so we can verify that the tracking is working.
//
// On failure, newGeneration returns nil values for generation and func,
// and an error. On success, the func returned is the close func to use
// when the generation is no longer needed by the caller.
func newGeneration(existing generation, path string, readData bool, setup func([]byte, *os.File, generation, bool) (bool, error), logger logger.Logger) (generation, error) {
m := mmapGeneration{path: path, logger: logger}
if existing != nil {
m.generation = existing.Generation() + 1
// we might keep a previous generation around just for its generation count.
if !existing.Dead() {
defer existing.Done()
}
}
shouldClose, err := m.openFile()
if err != nil {
return nil, err
}
m.id = fmt.Sprintf("%s:%d", m.path, m.generation)
// possibly assign new generation ID if this one's been used, which can
// happen with reopens, especially during testing.
m.id = registerGeneration(m.id)
// if debugging, we always want the finalizer on so we notice if a
// generation is finalized without being closed. for non-debugging
// use, we only need it when the generation is closed.
if generationDebug {
runtime.SetFinalizer(&m, generationFinalizer)
}
// Mmap the underlying file so it can be zero copied.
var mapped bool
var data []byte
fi, err := m.file.Stat()
if err == nil && fi.Size() > 0 {
data, err = syswrap.Mmap(int(m.file.Fd()), 0, int(fi.Size()), syscall.PROT_READ, syscall.MAP_SHARED)
if err == syswrap.ErrMaxMapCountReached {
// I have no idea where/how to display this message.
m.logger.Printf("maximum number of maps reached, reading file '%s' instead", m.path)
} else if err != nil {
m.Cancel()
return nil, errors.Wrap(err, "mmap failed")
} else {
mapped = true
}
}
if data == nil && readData {
data, err = ioutil.ReadAll(m.file)
if err != nil {
m.Cancel()
return nil, errors.Wrap(err, "failure file readall")
}
}
// if we got here, data's the expected data, so let's try to use it
mappedAny, err := setup(data, m.file, &m, mapped)
// if the setup failed, we unmap data if we previously mapped it,
// and exit. Note that having no data, or having only trivial
// data (like a zero-container Roaring file) isn't "failed".
if err != nil {
m.Cancel()
// Unless, that is, we think the file probably ought to
// be truncated: For instance, if a bitmap has a corrupted
// ops log, we could truncate that part of it and retry.
if err, ok := err.(roaring.FileShouldBeTruncatedError); ok && m.retries < 1 {
m.logger.Printf("file %s read partially, but should-be-truncated at %d bytes\n", m.path, err.SuggestedLength())
// close this generation, then try again. once.
m.retries++
err := os.Truncate(m.path, err.SuggestedLength())
if err != nil {
m.logger.Printf("truncating file failed [but retrying anyway]: %v\n", err)
}
return newGeneration(&m, path, readData, setup, logger)
}
return nil, err
}
if mapped {
// when generationDebug is on, we want to track this even
// if it's not being used.
if generationDebug || mappedAny {
// Advise the kernel that the mmap is accessed randomly.
// We don't care much about errors with this.
_ = madvise(data, syscall.MADV_RANDOM)
// store the data, so we can unmap it when this generation
// gets finalized.
m.data = data
} else {
// unmap the data and don't stash the pointer in this
// generation. It's not being used. This generation
// doesn't need to exist, yay.
unmapErr := syswrap.Munmap(data)
if unmapErr != nil {
m.logger.Printf("error unmapping (probably harmless): %v", unmapErr)
}
}
}
// shouldClose comes from underlying syswrap.OpenFile, which checks
// a count of open files to hint at us when we need to start closing
// files to preserve open file descriptor limit.
if shouldClose {
err := m.closeFile()
if err != nil {
m.logger.Printf("closing file to preserve open files failed: %v\n", err)
}
}
// It's possible that the generation has no actual data to track,
// because nothing's mapped, in which case there won't be any bitmap
// sources following this, just the fragment source. (Bitmaps won't
// be attached to the source unless they're actually mapped to it,
// or generationDebug is true). That's okay. We pay a tiny cost
// for the finalizer, but we also get higher confidence that it really
// does get cleaned up.
return &m, nil
}

160
generation_debug.go Normal file
View file

@ -0,0 +1,160 @@
// Copyright 2019 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
// +build generationdebug
package pilosa
import (
"fmt"
"math/rand"
"runtime"
"sort"
"sync"
"time"
)
const generationDebug = true
type lifespan struct {
from, to, finalized time.Time
}
var knownGenerations map[string]lifespan
var knownGenerationLock sync.Mutex
var timeZero time.Time
func registerGeneration(id string) string {
knownGenerationLock.Lock()
defer knownGenerationLock.Unlock()
if knownGenerations == nil {
knownGenerations = make(map[string]lifespan)
}
newSpan := lifespan{from: time.Now()}
origId := id
// if you have more than 65k of the same file open, maybe you have bigger
// problems than this.
for span, exists := knownGenerations[id]; exists; span, exists = knownGenerations[id] {
suffix := fmt.Sprintf("::%04x", rand.Int63n(65536))
if span.finalized != timeZero {
fmt.Printf("new generation %s: adding %s, previously existed, created %v, died %v, finalized %v\n",
id, suffix, span.from, span.to, span.finalized)
} else {
if span.to != timeZero {
fmt.Printf("new generation %s: adding %s, previously existed, created %v, died %v\n", id, suffix, span.from, span.to)
} else {
fmt.Printf("new generation %s: adding %s, already exists, created %v", id, suffix, span.from)
}
}
id = origId + suffix
}
fmt.Printf("new generation %s\n", id)
knownGenerations[id] = newSpan
return id
}
func endGeneration(id string) {
knownGenerationLock.Lock()
defer knownGenerationLock.Unlock()
span, exists := knownGenerations[id]
if !exists {
oops := fmt.Sprintf("ending generation %s: unknown", id)
panic(oops)
}
if span.finalized != timeZero || span.to != timeZero {
oops := fmt.Sprintf("ending generation %s: already died at %v, finalized at %v", id, span.to, span.finalized)
panic(oops)
}
span.to = time.Now()
knownGenerations[id] = span
}
// cancelGeneration marks the generation as finalized. In principle it's
// only used in cases where we just started a generation but something
// went wrong. it's not fancier than this because of the weird cases
// where the same generation shows up again, such as when closing and
// reopening an index so we don't know about previous instances of the
// same files.
func cancelGeneration(id string) {
knownGenerationLock.Lock()
defer knownGenerationLock.Unlock()
span, exists := knownGenerations[id]
if exists {
span.finalized = time.Now()
span.to = span.finalized
knownGenerations[id] = span
}
}
func finalizeGeneration(id string) {
knownGenerationLock.Lock()
defer knownGenerationLock.Unlock()
span, exists := knownGenerations[id]
if !exists {
oops := fmt.Sprintf("finalizing generation %s: unknown", id)
panic(oops)
}
if span.finalized != timeZero {
var oops string
if span.to != timeZero {
oops = fmt.Sprintf("finalizing generation %s: already finalized at %v, but not dead", id, span.finalized)
} else {
oops = fmt.Sprintf("finalizing generation %s: already finalized at %v, dead at %v", id, span.finalized, span.to)
}
panic(oops)
}
span.finalized = time.Now()
knownGenerations[id] = span
}
func reportGenerations() []string {
runtime.GC()
knownGenerationLock.Lock()
defer knownGenerationLock.Unlock()
var surviving []string
times := make([]int64, 0, len(knownGenerations))
for id, span := range knownGenerations {
if span.to == timeZero {
if span.finalized == timeZero {
surviving = append(surviving, fmt.Sprintf("%s: %v, not ended or finalized", id, span.from))
} else {
surviving = append(surviving, fmt.Sprintf("%s: %v, finalized %v, not ended", id, span.from, span.finalized))
}
} else {
if span.finalized == timeZero {
surviving = append(surviving, fmt.Sprintf("%s: %v to %v, not finalized", id, span.from, span.to))
} else {
times = append(times, int64(span.finalized.Sub(span.to)))
}
}
}
if len(times) > 0 {
sort.Slice(times, func(i, j int) bool { return times[i] < times[j] })
var total int64
for _, d := range times {
total += d
}
var mean, median, p90, p99, worst int64
mean = total / int64(len(times))
median = times[len(times)/2]
p90 = times[(len(times)*9)/10]
p99 = times[(len(times)*99)/100]
worst = times[len(times)-1]
surviving = append(surviving, fmt.Sprintf("%d finalized spans. lag: mean %v, median %v, p90 %v, p99 %v, worst %v",
len(times), time.Duration(mean), time.Duration(median), time.Duration(p90), time.Duration(p99), time.Duration(worst)))
}
return surviving
}

37
generation_nodebug.go Normal file
View file

@ -0,0 +1,37 @@
// Copyright 2019 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
// +build !generationdebug
package pilosa
const generationDebug = false
func registerGeneration(id string) string {
return id
}
func endGeneration(id string) {
}
func cancelGeneration(id string) {
}
func finalizeGeneration(id string) {
}
//lint:ignore U1000 this is conditional on a build flag, see generation_test.go.
func reportGenerations() []string { //nolint:unused,deadcode
return nil
}

39
generation_test.go Normal file
View file

@ -0,0 +1,39 @@
// Copyright 2019 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//
// +build generationdebug
package pilosa
import (
"fmt"
"os"
"testing"
)
func examineResults() {
results := reportGenerations()
if len(results) > 0 {
fmt.Printf("generations:\n")
for _, res := range results {
fmt.Printf(" %s\n", res)
}
}
}
func TestMain(m *testing.M) {
ret := m.Run()
examineResults()
os.Exit(ret)
}

12
go.mod
View file

@ -3,7 +3,6 @@ module github.com/pilosa/pilosa/v2
replace github.com/hashicorp/memberlist => github.com/pilosa/memberlist v0.1.4-0.20190415211605-f6512523c021 replace github.com/hashicorp/memberlist => github.com/pilosa/memberlist v0.1.4-0.20190415211605-f6512523c021
require ( require (
github.com/BurntSushi/toml v0.3.1 // indirect
github.com/CAFxX/gcnotifier v0.0.0-20190112062741-224a280d589d github.com/CAFxX/gcnotifier v0.0.0-20190112062741-224a280d589d
github.com/DataDog/datadog-go v0.0.0-20180822151419-281ae9f2d895 github.com/DataDog/datadog-go v0.0.0-20180822151419-281ae9f2d895
github.com/StackExchange/wmi v0.0.0-20190523213315-cbe66965904d // indirect github.com/StackExchange/wmi v0.0.0-20190523213315-cbe66965904d // indirect
@ -13,12 +12,14 @@ require (
github.com/davecgh/go-spew v1.1.1 github.com/davecgh/go-spew v1.1.1
github.com/go-ole/go-ole v1.2.4 // indirect github.com/go-ole/go-ole v1.2.4 // indirect
github.com/gogo/protobuf v1.2.0 github.com/gogo/protobuf v1.2.0
github.com/golang/protobuf v1.3.1 github.com/golang/protobuf v1.3.2
github.com/google/go-cmp v0.2.0 github.com/google/go-cmp v0.2.0
github.com/gorilla/handlers v1.3.0 github.com/gorilla/handlers v1.3.0
github.com/gorilla/mux v1.7.0 github.com/gorilla/mux v1.7.0
github.com/hashicorp/memberlist v0.1.3 github.com/hashicorp/memberlist v0.1.3
github.com/inconshreveable/mousetrap v1.0.0 // indirect github.com/inconshreveable/mousetrap v1.0.0 // indirect
github.com/molecula/ext v0.0.0-20200103203257-8a458a73e8c2
github.com/molecula/extensions v0.0.0-20191218165536-562244600fd4
github.com/opentracing/opentracing-go v1.1.0 github.com/opentracing/opentracing-go v1.1.0
github.com/pelletier/go-toml v1.2.0 github.com/pelletier/go-toml v1.2.0
github.com/pkg/errors v0.8.1 github.com/pkg/errors v0.8.1
@ -34,14 +35,17 @@ require (
github.com/uber-go/atomic v1.4.0 // indirect github.com/uber-go/atomic v1.4.0 // indirect
github.com/uber/jaeger-client-go v2.16.0+incompatible github.com/uber/jaeger-client-go v2.16.0+incompatible
github.com/uber/jaeger-lib v2.2.0+incompatible // indirect github.com/uber/jaeger-lib v2.2.0+incompatible // indirect
github.com/youtube/vitess v2.1.1+incompatible // indirect
go.uber.org/atomic v1.4.0 // indirect go.uber.org/atomic v1.4.0 // indirect
golang.org/x/crypto v0.0.0-20190426145343-a29dc8fdc734 // indirect golang.org/x/crypto v0.0.0-20190426145343-a29dc8fdc734 // indirect
golang.org/x/net v0.0.0-20190424112056-4829fb13d2c6 // indirect golang.org/x/net v0.0.0-20190424112056-4829fb13d2c6
golang.org/x/sync v0.0.0-20190423024810-112230192c58 golang.org/x/sync v0.0.0-20190423024810-112230192c58
golang.org/x/sys v0.0.0-20190429190828-d89cdac9e872 // indirect golang.org/x/sys v0.0.0-20190429190828-d89cdac9e872 // indirect
golang.org/x/text v0.3.2 // indirect golang.org/x/text v0.3.2 // indirect
google.golang.org/grpc v1.24.0
modernc.org/mathutil v1.0.0 modernc.org/mathutil v1.0.0
modernc.org/strutil v1.0.0 modernc.org/strutil v1.0.0
vitess.io/vitess v2.1.1+incompatible // indirect
) )
go 1.11 go 1.13

30
go.sum
View file

@ -1,3 +1,4 @@
cloud.google.com/go v0.26.0/go.mod h1:aQUYkXzVsufM+DwF1aE+0xfcU+56JwCaLick0ClmMTw=
github.com/BurntSushi/toml v0.3.1 h1:WXkYYl6Yr3qBf1K79EBnL4mak0OimBfB0XUf9Vl28OQ= github.com/BurntSushi/toml v0.3.1 h1:WXkYYl6Yr3qBf1K79EBnL4mak0OimBfB0XUf9Vl28OQ=
github.com/BurntSushi/toml v0.3.1/go.mod h1:xHWCNGjB5oqiDr8zfno3MHue2Ht5sIBksp03qcyfWMU= github.com/BurntSushi/toml v0.3.1/go.mod h1:xHWCNGjB5oqiDr8zfno3MHue2Ht5sIBksp03qcyfWMU=
github.com/CAFxX/gcnotifier v0.0.0-20190112062741-224a280d589d h1:n0G4ckjMEj7bWuGYUX0i8YlBeBBJuZ+HEHvHfyBDZtI= github.com/CAFxX/gcnotifier v0.0.0-20190112062741-224a280d589d h1:n0G4ckjMEj7bWuGYUX0i8YlBeBBJuZ+HEHvHfyBDZtI=
@ -20,6 +21,7 @@ github.com/boltdb/bolt v1.3.1 h1:JQmyP4ZBrce+ZQu0dY660FMfatumYDLun9hBCUVIkF4=
github.com/boltdb/bolt v1.3.1/go.mod h1:clJnj/oiGkjum5o1McbSZDSLxVThjynRyGBgiAx27Ps= github.com/boltdb/bolt v1.3.1/go.mod h1:clJnj/oiGkjum5o1McbSZDSLxVThjynRyGBgiAx27Ps=
github.com/cespare/xxhash v1.1.0 h1:a6HrQnmkObjyL+Gs60czilIUGqrzKutQD6XZog3p+ko= github.com/cespare/xxhash v1.1.0 h1:a6HrQnmkObjyL+Gs60czilIUGqrzKutQD6XZog3p+ko=
github.com/cespare/xxhash v1.1.0/go.mod h1:XrSqR1VqqWfGrhpAt58auRo0WTKS1nRRg3ghfAqPWnc= github.com/cespare/xxhash v1.1.0/go.mod h1:XrSqR1VqqWfGrhpAt58auRo0WTKS1nRRg3ghfAqPWnc=
github.com/client9/misspell v0.3.4/go.mod h1:qj6jICC3Q7zFZvVWo7KLAzC3yx5G7kyvSDkc90ppPyw=
github.com/codahale/hdrhistogram v0.0.0-20161010025455-3a0bb77429bd h1:qMd81Ts1T2OTKmB4acZcyKaMtRnY5Y44NuXGX2GFJ1w= github.com/codahale/hdrhistogram v0.0.0-20161010025455-3a0bb77429bd h1:qMd81Ts1T2OTKmB4acZcyKaMtRnY5Y44NuXGX2GFJ1w=
github.com/codahale/hdrhistogram v0.0.0-20161010025455-3a0bb77429bd/go.mod h1:sE/e/2PUdi/liOCUjSTXgM1o87ZssimdTWN964YiIeI= github.com/codahale/hdrhistogram v0.0.0-20161010025455-3a0bb77429bd/go.mod h1:sE/e/2PUdi/liOCUjSTXgM1o87ZssimdTWN964YiIeI=
github.com/coreos/etcd v3.3.10+incompatible/go.mod h1:uF7uidLiAD3TWHmW31ZFd/JWoc32PjwdhPthX9715RE= github.com/coreos/etcd v3.3.10+incompatible/go.mod h1:uF7uidLiAD3TWHmW31ZFd/JWoc32PjwdhPthX9715RE=
@ -39,10 +41,15 @@ github.com/go-stack/stack v1.8.0/go.mod h1:v0f6uXyyMGvRgIKkXu+yp6POWl0qKG85gN/me
github.com/gogo/protobuf v1.1.1/go.mod h1:r8qH/GZQm5c6nD/R0oafs1akxWv10x8SbQlK7atdtwQ= github.com/gogo/protobuf v1.1.1/go.mod h1:r8qH/GZQm5c6nD/R0oafs1akxWv10x8SbQlK7atdtwQ=
github.com/gogo/protobuf v1.2.0 h1:xU6/SpYbvkNYiptHJYEDRseDLvYE7wSqhYYNy0QSUzI= github.com/gogo/protobuf v1.2.0 h1:xU6/SpYbvkNYiptHJYEDRseDLvYE7wSqhYYNy0QSUzI=
github.com/gogo/protobuf v1.2.0/go.mod h1:r8qH/GZQm5c6nD/R0oafs1akxWv10x8SbQlK7atdtwQ= github.com/gogo/protobuf v1.2.0/go.mod h1:r8qH/GZQm5c6nD/R0oafs1akxWv10x8SbQlK7atdtwQ=
github.com/golang/glog v0.0.0-20160126235308-23def4e6c14b h1:VKtxabqXZkF25pY9ekfRL6a582T4P37/31XEstQ5p58=
github.com/golang/glog v0.0.0-20160126235308-23def4e6c14b/go.mod h1:SBH7ygxi8pfUlaOkMMuAQtPIUF8ecWP5IEl/CR7VP2Q=
github.com/golang/mock v1.1.1/go.mod h1:oTYuIxOrZwtPieC+H1uAHpcLFnEyAGVDL/k47Jfbm0A=
github.com/golang/protobuf v1.2.0 h1:P3YflyNX/ehuJFLhxviNdFxQPkGK5cDcApsge1SqnvM= github.com/golang/protobuf v1.2.0 h1:P3YflyNX/ehuJFLhxviNdFxQPkGK5cDcApsge1SqnvM=
github.com/golang/protobuf v1.2.0/go.mod h1:6lQm79b+lXiMfvg/cZm0SGofjICqVBUtrP5yJMmIC1U= github.com/golang/protobuf v1.2.0/go.mod h1:6lQm79b+lXiMfvg/cZm0SGofjICqVBUtrP5yJMmIC1U=
github.com/golang/protobuf v1.3.1 h1:YF8+flBXS5eO826T4nzqPrxfhQThhXl0YzfuUPu4SBg= github.com/golang/protobuf v1.3.1 h1:YF8+flBXS5eO826T4nzqPrxfhQThhXl0YzfuUPu4SBg=
github.com/golang/protobuf v1.3.1/go.mod h1:6lQm79b+lXiMfvg/cZm0SGofjICqVBUtrP5yJMmIC1U= github.com/golang/protobuf v1.3.1/go.mod h1:6lQm79b+lXiMfvg/cZm0SGofjICqVBUtrP5yJMmIC1U=
github.com/golang/protobuf v1.3.2 h1:6nsPYzhq5kReh6QImI3k5qWzO4PEbvbIW2cwSfR/6xs=
github.com/golang/protobuf v1.3.2/go.mod h1:6lQm79b+lXiMfvg/cZm0SGofjICqVBUtrP5yJMmIC1U=
github.com/google/btree v0.0.0-20180813153112-4030bb1f1f0c h1:964Od4U6p2jUkFxvCydnIczKteheJEzHRToSGK3Bnlw= github.com/google/btree v0.0.0-20180813153112-4030bb1f1f0c h1:964Od4U6p2jUkFxvCydnIczKteheJEzHRToSGK3Bnlw=
github.com/google/btree v0.0.0-20180813153112-4030bb1f1f0c/go.mod h1:lNA+9X1NB3Zf8V7Ke586lFgjr2dZNuvo3lPJSGZ5JPQ= github.com/google/btree v0.0.0-20180813153112-4030bb1f1f0c/go.mod h1:lNA+9X1NB3Zf8V7Ke586lFgjr2dZNuvo3lPJSGZ5JPQ=
github.com/google/go-cmp v0.2.0 h1:+dTQ8DZQJz0Mb/HjFlkptS1FeQ4cWSnN941F8aEG4SQ= github.com/google/go-cmp v0.2.0 h1:+dTQ8DZQJz0Mb/HjFlkptS1FeQ4cWSnN941F8aEG4SQ=
@ -80,6 +87,14 @@ github.com/miekg/dns v1.0.14 h1:9jZdLNd/P4+SfEJ0TNyxYpsK8N4GtfylBLqtbYN1sbA=
github.com/miekg/dns v1.0.14/go.mod h1:W1PPwlIAgtquWBMBEV9nkV9Cazfe8ScdGz/Lj7v3Nrg= github.com/miekg/dns v1.0.14/go.mod h1:W1PPwlIAgtquWBMBEV9nkV9Cazfe8ScdGz/Lj7v3Nrg=
github.com/mitchellh/mapstructure v1.1.2 h1:fmNYVwqnSfB9mZU6OS2O6GsXM+wcskZDuKQzvN1EDeE= github.com/mitchellh/mapstructure v1.1.2 h1:fmNYVwqnSfB9mZU6OS2O6GsXM+wcskZDuKQzvN1EDeE=
github.com/mitchellh/mapstructure v1.1.2/go.mod h1:FVVH3fgwuzCH5S8UJGiWEs2h04kUh9fWfEaFds41c1Y= github.com/mitchellh/mapstructure v1.1.2/go.mod h1:FVVH3fgwuzCH5S8UJGiWEs2h04kUh9fWfEaFds41c1Y=
github.com/molecula/apophenia v0.0.0-20190827192002-68b7a14a478b h1:cZADDaNYM7xn/nklO3g198JerGQjadFuA0ofxBJgK0Y=
github.com/molecula/apophenia v0.0.0-20190827192002-68b7a14a478b/go.mod h1:uXd1BiH7xLmgkhVmspdJLENv6uGWrTL/MQX2TN7Yz9s=
github.com/molecula/ext v0.0.0-20191202195653-240f38a75171 h1:4VK7u/RM+54Yaz8aRB9vIaDSnbKi3M0NQYg5tsZvOT4=
github.com/molecula/ext v0.0.0-20191202195653-240f38a75171/go.mod h1:r6EIj0GH8dx5xxFLW6Voi1/mX3wXOUkJu6AoEE/xvGQ=
github.com/molecula/ext v0.0.0-20200103203257-8a458a73e8c2 h1:XOImsA5XhGklFj8Y0TxSm1qWZzEwYxom2JOXiu9GMq0=
github.com/molecula/ext v0.0.0-20200103203257-8a458a73e8c2/go.mod h1:r6EIj0GH8dx5xxFLW6Voi1/mX3wXOUkJu6AoEE/xvGQ=
github.com/molecula/extensions v0.0.0-20191218165536-562244600fd4 h1:mDB/dicofRVFuRYcCVPk+JBiVKXlfbzMahuqHvrYqu4=
github.com/molecula/extensions v0.0.0-20191218165536-562244600fd4/go.mod h1:QQgN5OFjuBAi4Q2UYVMzfvi4k9yvg/qqC+MNFB4I9JI=
github.com/mwitkow/go-conntrack v0.0.0-20161129095857-cc309e4a2223/go.mod h1:qRWi+5nqEBWmkhHvq77mSJWrCKwh8bxhgT7d/eI7P4U= github.com/mwitkow/go-conntrack v0.0.0-20161129095857-cc309e4a2223/go.mod h1:qRWi+5nqEBWmkhHvq77mSJWrCKwh8bxhgT7d/eI7P4U=
github.com/oklog/ulid v1.3.1/go.mod h1:CirwcVhetQ6Lv90oh/F+FBtV6XMibvdAFo93nm5qn4U= github.com/oklog/ulid v1.3.1/go.mod h1:CirwcVhetQ6Lv90oh/F+FBtV6XMibvdAFo93nm5qn4U=
github.com/opentracing/opentracing-go v1.1.0 h1:pWlfV3Bxv7k65HYwkikxat0+s3pV4bsqf19k25Ur8rU= github.com/opentracing/opentracing-go v1.1.0 h1:pWlfV3Bxv7k65HYwkikxat0+s3pV4bsqf19k25Ur8rU=
@ -144,6 +159,8 @@ github.com/uber/jaeger-lib v2.2.0+incompatible h1:MxZXOiR2JuoANZ3J6DE/U0kSFv/eJ/
github.com/uber/jaeger-lib v2.2.0+incompatible/go.mod h1:ComeNDZlWwrWnDv8aPp0Ba6+uUTzImX/AauajbLI56U= github.com/uber/jaeger-lib v2.2.0+incompatible/go.mod h1:ComeNDZlWwrWnDv8aPp0Ba6+uUTzImX/AauajbLI56U=
github.com/ugorji/go/codec v0.0.0-20181204163529-d75b2dcb6bc8/go.mod h1:VFNgLljTbGfSG7qAOspJ7OScBnGdDN/yBr0sguwnwf0= github.com/ugorji/go/codec v0.0.0-20181204163529-d75b2dcb6bc8/go.mod h1:VFNgLljTbGfSG7qAOspJ7OScBnGdDN/yBr0sguwnwf0=
github.com/xordataexchange/crypt v0.0.3-0.20170626215501-b2862e3d0a77/go.mod h1:aYKd//L2LvnjZzWKhF00oedf4jCCReLcmhLdhm1A27Q= github.com/xordataexchange/crypt v0.0.3-0.20170626215501-b2862e3d0a77/go.mod h1:aYKd//L2LvnjZzWKhF00oedf4jCCReLcmhLdhm1A27Q=
github.com/youtube/vitess v2.1.1+incompatible h1:SE+P7DNX/jw5RHFs5CHRhZQjq402EJFCD33JhzQMdDw=
github.com/youtube/vitess v2.1.1+incompatible/go.mod h1:hpMim5/30F1r+0P8GGtB29d0gWHr0IZ5unS+CG0zMx8=
go.uber.org/atomic v1.4.0 h1:cxzIVoETapQEqDhQu3QfnvXAV4AlzcvUCxkVUFw3+EU= go.uber.org/atomic v1.4.0 h1:cxzIVoETapQEqDhQu3QfnvXAV4AlzcvUCxkVUFw3+EU=
go.uber.org/atomic v1.4.0/go.mod h1:gD2HeocX3+yG+ygLZcrzQJaqmWj9AIm7n08wl/qW/PE= go.uber.org/atomic v1.4.0/go.mod h1:gD2HeocX3+yG+ygLZcrzQJaqmWj9AIm7n08wl/qW/PE=
golang.org/x/crypto v0.0.0-20180904163835-0709b304e793/go.mod h1:6SG95UA2DQfeDnfUPMdvaQW0Q7yPrPDi9nlGo2tz2b4= golang.org/x/crypto v0.0.0-20180904163835-0709b304e793/go.mod h1:6SG95UA2DQfeDnfUPMdvaQW0Q7yPrPDi9nlGo2tz2b4=
@ -153,13 +170,16 @@ golang.org/x/crypto v0.0.0-20181203042331-505ab145d0a9/go.mod h1:6SG95UA2DQfeDnf
golang.org/x/crypto v0.0.0-20190308221718-c2843e01d9a2/go.mod h1:djNgcEr1/C05ACkg1iLfiJU5Ep61QUkGW8qpdssI0+w= golang.org/x/crypto v0.0.0-20190308221718-c2843e01d9a2/go.mod h1:djNgcEr1/C05ACkg1iLfiJU5Ep61QUkGW8qpdssI0+w=
golang.org/x/crypto v0.0.0-20190426145343-a29dc8fdc734 h1:p/H982KKEjUnLJkM3tt/LemDnOc1GiZL5FCVlORJ5zo= golang.org/x/crypto v0.0.0-20190426145343-a29dc8fdc734 h1:p/H982KKEjUnLJkM3tt/LemDnOc1GiZL5FCVlORJ5zo=
golang.org/x/crypto v0.0.0-20190426145343-a29dc8fdc734/go.mod h1:yigFU9vqHzYiE8UmvKecakEJjdnWj3jj499lnFckfCI= golang.org/x/crypto v0.0.0-20190426145343-a29dc8fdc734/go.mod h1:yigFU9vqHzYiE8UmvKecakEJjdnWj3jj499lnFckfCI=
golang.org/x/lint v0.0.0-20190313153728-d0100b6bd8b3/go.mod h1:6SW0HCj/g11FgYtHlgUYUwCkIfeOF89ocIRzGO/8vkc=
golang.org/x/net v0.0.0-20181023162649-9b4f9f5ad519 h1:x6rhz8Y9CjbgQkccRGmELH6K+LJj7tOoh3XWeC1yaQM= golang.org/x/net v0.0.0-20181023162649-9b4f9f5ad519 h1:x6rhz8Y9CjbgQkccRGmELH6K+LJj7tOoh3XWeC1yaQM=
golang.org/x/net v0.0.0-20181023162649-9b4f9f5ad519/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4= golang.org/x/net v0.0.0-20181023162649-9b4f9f5ad519/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20181114220301-adae6a3d119a h1:gOpx8G595UYyvj8UK4+OFyY4rx037g3fmfhe5SasG3U= golang.org/x/net v0.0.0-20181114220301-adae6a3d119a h1:gOpx8G595UYyvj8UK4+OFyY4rx037g3fmfhe5SasG3U=
golang.org/x/net v0.0.0-20181114220301-adae6a3d119a/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4= golang.org/x/net v0.0.0-20181114220301-adae6a3d119a/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20190311183353-d8887717615a/go.mod h1:t9HGtf8HONx5eT2rtn7q6eTqICYqUVnKs3thJo3Qplg=
golang.org/x/net v0.0.0-20190404232315-eb5bcb51f2a3/go.mod h1:t9HGtf8HONx5eT2rtn7q6eTqICYqUVnKs3thJo3Qplg= golang.org/x/net v0.0.0-20190404232315-eb5bcb51f2a3/go.mod h1:t9HGtf8HONx5eT2rtn7q6eTqICYqUVnKs3thJo3Qplg=
golang.org/x/net v0.0.0-20190424112056-4829fb13d2c6 h1:FP8hkuE6yUEaJnK7O2eTuejKWwW+Rhfj80dQ2JcKxCU= golang.org/x/net v0.0.0-20190424112056-4829fb13d2c6 h1:FP8hkuE6yUEaJnK7O2eTuejKWwW+Rhfj80dQ2JcKxCU=
golang.org/x/net v0.0.0-20190424112056-4829fb13d2c6/go.mod h1:t9HGtf8HONx5eT2rtn7q6eTqICYqUVnKs3thJo3Qplg= golang.org/x/net v0.0.0-20190424112056-4829fb13d2c6/go.mod h1:t9HGtf8HONx5eT2rtn7q6eTqICYqUVnKs3thJo3Qplg=
golang.org/x/oauth2 v0.0.0-20180821212333-d2e6202438be/go.mod h1:N/0e6XlmueqKjAGxoOufVs8QHGRruUQn6yWY3a++T0U=
golang.org/x/sync v0.0.0-20181108010431-42b317875d0f/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM= golang.org/x/sync v0.0.0-20181108010431-42b317875d0f/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20181221193216-37e7f081c4d4 h1:YUO/7uOKsKeq9UokNS62b8FYywz3ker1l1vDZRCRefw= golang.org/x/sync v0.0.0-20181221193216-37e7f081c4d4 h1:YUO/7uOKsKeq9UokNS62b8FYywz3ker1l1vDZRCRefw=
golang.org/x/sync v0.0.0-20181221193216-37e7f081c4d4/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM= golang.org/x/sync v0.0.0-20181221193216-37e7f081c4d4/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
@ -180,13 +200,23 @@ golang.org/x/text v0.3.0/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.2 h1:tW2bmiBqwgJj/UpqtC8EpXEZVYOwU0yG4iWbprSVAcs= golang.org/x/text v0.3.2 h1:tW2bmiBqwgJj/UpqtC8EpXEZVYOwU0yG4iWbprSVAcs=
golang.org/x/text v0.3.2/go.mod h1:bEr9sfX3Q8Zfm5fL9x+3itogRgK3+ptLWKqgva+5dAk= golang.org/x/text v0.3.2/go.mod h1:bEr9sfX3Q8Zfm5fL9x+3itogRgK3+ptLWKqgva+5dAk=
golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ= golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ=
golang.org/x/tools v0.0.0-20190311212946-11955173bddd/go.mod h1:LCzVGOaR6xXOjkQ3onu1FJEFr0SW1gC7cKk1uF8kGRs=
golang.org/x/tools v0.0.0-20190524140312-2c0ae7006135/go.mod h1:RgjU9mgBXZiqYHBnxXauZ1Gv1EHHAz9KjViQ78xBX0Q=
google.golang.org/appengine v1.1.0/go.mod h1:EbEs0AVv82hx2wNQdGPgUI5lhzA/G0D9YwlJXL52JkM=
google.golang.org/genproto v0.0.0-20180817151627-c66870c02cf8 h1:Nw54tB0rB7hY/N0NQvRW8DG4Yk3Q6T9cu9RcFQDu1tc=
google.golang.org/genproto v0.0.0-20180817151627-c66870c02cf8/go.mod h1:JiN7NxoALGmiZfu7CAH4rXhgtRTLTxftemlI0sWmxmc=
google.golang.org/grpc v1.24.0 h1:vb/1TCsVn3DcJlQ0Gs1yB1pKI6Do2/QNwxdKqmc/b0s=
google.golang.org/grpc v1.24.0/go.mod h1:XDChyiUovWa60DnaeDeZmSW86xtLtjtZbwvSiRnRtcA=
gopkg.in/alecthomas/kingpin.v2 v2.2.6/go.mod h1:FMv+mEhP44yOT+4EoQTLFTRgOQ1FBLkstjWtayDeSgw= gopkg.in/alecthomas/kingpin.v2 v2.2.6/go.mod h1:FMv+mEhP44yOT+4EoQTLFTRgOQ1FBLkstjWtayDeSgw=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405 h1:yhCVgyC4o1eVCa2tZl7eS0r+SDo693bJlVdllGtEeKM= gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405 h1:yhCVgyC4o1eVCa2tZl7eS0r+SDo693bJlVdllGtEeKM=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0= gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/yaml.v2 v2.2.1/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI= gopkg.in/yaml.v2 v2.2.1/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI=
gopkg.in/yaml.v2 v2.2.2 h1:ZCJp+EgiOT7lHqUV2J862kp8Qj64Jo6az82+3Td9dZw= gopkg.in/yaml.v2 v2.2.2 h1:ZCJp+EgiOT7lHqUV2J862kp8Qj64Jo6az82+3Td9dZw=
gopkg.in/yaml.v2 v2.2.2/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI= gopkg.in/yaml.v2 v2.2.2/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI=
honnef.co/go/tools v0.0.0-20190523083050-ea95bdfd59fc/go.mod h1:rf3lG4BRIbNafJWhAfAdb/ePZxsR/4RtNHQocxwk9r4=
modernc.org/mathutil v1.0.0 h1:93vKjrJopTPrtTNpZ8XIovER7iCIH1QU7wNbOQXC60I= modernc.org/mathutil v1.0.0 h1:93vKjrJopTPrtTNpZ8XIovER7iCIH1QU7wNbOQXC60I=
modernc.org/mathutil v1.0.0/go.mod h1:wU0vUrJsVWBZ4P6e7xtFJEhFSNsfRLJ8H458uRjg03k= modernc.org/mathutil v1.0.0/go.mod h1:wU0vUrJsVWBZ4P6e7xtFJEhFSNsfRLJ8H458uRjg03k=
modernc.org/strutil v1.0.0 h1:XVFtQwFVwc02Wk+0L/Z/zDDXO81r5Lhe6iMKmGX3KhE= modernc.org/strutil v1.0.0 h1:XVFtQwFVwc02Wk+0L/Z/zDDXO81r5Lhe6iMKmGX3KhE=
modernc.org/strutil v1.0.0/go.mod h1:lstksw84oURvj9y3tn8lGvRxyRC1S2+g5uuIzNfIOBs= modernc.org/strutil v1.0.0/go.mod h1:lstksw84oURvj9y3tn8lGvRxyRC1S2+g5uuIzNfIOBs=
vitess.io/vitess v2.1.1+incompatible h1:nuuGHiWYWpudD3gOCLeGzol2EJ25e/u5Wer2wV1O130=
vitess.io/vitess v2.1.1+incompatible/go.mod h1:h4qvkyNYTOC0xI+vcidSWoka0gQAZc9ZPHbkHo48gP0=

View file

@ -16,6 +16,9 @@ package pilosa
import ( import (
"encoding/json" "encoding/json"
"github.com/pilosa/pilosa/v2/tracing"
"github.com/pkg/errors"
) )
// QueryRequest represent a request to process a query. // QueryRequest represent a request to process a query.
@ -42,6 +45,13 @@ type QueryRequest struct {
// If true, indicates that query is part of a larger distributed query. // If true, indicates that query is part of a larger distributed query.
// If false, this request is on the originating node. // If false, this request is on the originating node.
Remote bool Remote bool
// Should we profile this query?
Profile bool
// Additional data associated with the query, in cases where there's
// row-style inputs for precomputed values.
EmbeddedData []*Row
} }
// QueryResponse represent a response from a processed query. // QueryResponse represent a response from a processed query.
@ -55,6 +65,9 @@ type QueryResponse struct {
// Error during parsing or execution. // Error during parsing or execution.
Err error Err error
// Profiling data, if any
Profile *tracing.Profile
} }
// MarshalJSON marshals QueryResponse into a JSON-encoded byte slice // MarshalJSON marshals QueryResponse into a JSON-encoded byte slice
@ -68,9 +81,11 @@ func (resp *QueryResponse) MarshalJSON() ([]byte, error) {
return json.Marshal(struct { return json.Marshal(struct {
Results []interface{} `json:"results"` Results []interface{} `json:"results"`
ColumnAttrSets []*ColumnAttrSet `json:"columnAttrs,omitempty"` ColumnAttrSets []*ColumnAttrSet `json:"columnAttrs,omitempty"`
Profile *tracing.Profile `json:"profile,omitempty"`
}{ }{
Results: resp.Results, Results: resp.Results,
ColumnAttrSets: resp.ColumnAttrSets, ColumnAttrSets: resp.ColumnAttrSets,
Profile: resp.Profile,
}) })
} }
@ -99,10 +114,61 @@ var NopHandler Handler = nopHandler{}
type ImportValueRequest struct { type ImportValueRequest struct {
Index string Index string
Field string Field string
// if Shard is MaxUint64 (an impossible shard value), this
// indicates that the column IDs may come from multiple shards.
Shard uint64 Shard uint64
ColumnIDs []uint64 ColumnIDs []uint64
ColumnKeys []string ColumnKeys []string
Values []int64 Values []int64
FloatValues []float64
StringValues []string
}
func (ivr *ImportValueRequest) Len() int { return len(ivr.ColumnIDs) }
func (ivr *ImportValueRequest) Less(i, j int) bool { return ivr.ColumnIDs[i] < ivr.ColumnIDs[j] }
func (ivr *ImportValueRequest) Swap(i, j int) {
ivr.ColumnIDs[i], ivr.ColumnIDs[j] = ivr.ColumnIDs[j], ivr.ColumnIDs[i]
if len(ivr.Values) > 0 {
ivr.Values[i], ivr.Values[j] = ivr.Values[j], ivr.Values[i]
} else if len(ivr.FloatValues) > 0 {
ivr.FloatValues[i], ivr.FloatValues[j] = ivr.FloatValues[j], ivr.FloatValues[i]
} else if len(ivr.StringValues) > 0 {
ivr.StringValues[i], ivr.StringValues[j] = ivr.StringValues[j], ivr.StringValues[i]
}
}
// Validate ensures that the payload of the request is valid.
func (ivr *ImportValueRequest) Validate() error {
if ivr.Index == "" || ivr.Field == "" {
return errors.Errorf("index and field required, but got '%s' and '%s'", ivr.Index, ivr.Field)
}
if len(ivr.ColumnIDs) != 0 && len(ivr.ColumnKeys) != 0 {
return errors.Errorf("must pass either column ids or keys, but not both")
}
var valueSetCount int
if len(ivr.Values) != 0 {
valueSetCount++
}
if len(ivr.FloatValues) != 0 {
valueSetCount++
}
if len(ivr.StringValues) != 0 {
valueSetCount++
}
if valueSetCount > 1 {
return errors.Errorf("must pass ints, floats, or strings but not multiple")
}
return nil
}
// ImportColumnAttrsRequest describes the import request structure
// for a ColumnAttr import
type ImportColumnAttrsRequest struct {
AttrKey string
ColumnIDs []uint64
AttrVals []string
Shard int64
Index string
} }
// ImportRequest describes the import request structure // ImportRequest describes the import request structure

View file

@ -75,7 +75,7 @@ type Holder struct {
Logger logger.Logger Logger logger.Logger
snapshotQueue chan *fragment snapshotQueue snapshotQueue
// Manages replication from the primary node. // Manages replication from the primary node.
primaryTranslateNode *Node primaryTranslateNode *Node
@ -84,6 +84,16 @@ type Holder struct {
// Instantiates new translation stores for indexes & fields. // Instantiates new translation stores for indexes & fields.
OpenTranslateStore OpenTranslateStoreFunc // local store OpenTranslateStore OpenTranslateStoreFunc // local store
OpenTranslateReader OpenTranslateReaderFunc // replication OpenTranslateReader OpenTranslateReaderFunc // replication
// Queue of fields (having a foreign index) which have
// opened before their foreign index has opened.
foreignIndexFields []*Field
// opening is set to true while Holder is opening.
// It's used to determine if foreign index application
// needs to be queued and completed after all indexes
// have opened.
opening bool
} }
// lockedChan looks a little ridiculous admittedly, but exists for good reason. // lockedChan looks a little ridiculous admittedly, but exists for good reason.
@ -135,6 +145,9 @@ func NewHolder() *Holder {
// Open initializes the root data directory for the holder. // Open initializes the root data directory for the holder.
func (h *Holder) Open() error { func (h *Holder) Open() error {
h.opening = true
defer func() { h.opening = false }()
// Reset closing in case Holder is being reopened. // Reset closing in case Holder is being reopened.
h.closing = make(chan struct{}) h.closing = make(chan struct{})
@ -167,7 +180,7 @@ func (h *Holder) Open() error {
// Run snapshots asynchronously. The snapshotQueue will have a background // Run snapshots asynchronously. The snapshotQueue will have a background
// task associated with it which flushes it and waits until this channel // task associated with it which flushes it and waits until this channel
// is closed, so we should always close this channel when done. // is closed, so we should always close this channel when done.
h.snapshotQueue = newSnapshotQueue(100, 2, h.Logger) h.snapshotQueue = newSnapshotQueue(10, 2, h.Logger)
for _, fi := range fis { for _, fi := range fis {
// Skip files or hidden directories. // Skip files or hidden directories.
@ -196,6 +209,14 @@ func (h *Holder) Open() error {
h.indexes[index.Name()] = index h.indexes[index.Name()] = index
h.mu.Unlock() h.mu.Unlock()
} }
// If any fields were opened before their foreign index
// was opened, it's safe to process those now since all index
// opens have completed by this point.
if err := h.processForeignIndexFields(); err != nil {
return errors.Wrap(err, "processing foreign index fields")
}
h.Logger.Printf("open holder: complete") h.Logger.Printf("open holder: complete")
// Periodically flush cache. // Periodically flush cache.
@ -203,11 +224,39 @@ func (h *Holder) Open() error {
go func() { defer h.wg.Done(); h.monitorCacheFlush() }() go func() { defer h.wg.Done(); h.monitorCacheFlush() }()
h.Stats.Open() h.Stats.Open()
h.snapshotQueue.ScanHolder(h)
h.opened.Close() h.opened.Close()
return nil return nil
} }
// checkForeignIndex is a check before applying a foreign
// index to a field; if the index is not yet available,
// (because holder is still opening and may not have opened
// the index yet), this method queues it up to be processed
// once all indexes have been opened.
func (h *Holder) checkForeignIndex(f *Field) error {
if h.opening {
if fi := h.Index(f.options.ForeignIndex); fi == nil {
h.foreignIndexFields = append(h.foreignIndexFields, f)
return nil
}
}
return f.applyForeignIndex()
}
// processForeignIndexFields applies a foreign index to any
// fields which were opened before their foreign index.
func (h *Holder) processForeignIndexFields() error {
for _, f := range h.foreignIndexFields {
if err := f.applyForeignIndex(); err != nil {
return errors.Wrap(err, "applying foreign index")
}
}
h.foreignIndexFields = h.foreignIndexFields[:0] // reset
return nil
}
// Close closes all open fragments. // Close closes all open fragments.
func (h *Holder) Close() error { func (h *Holder) Close() error {
h.Stats.Close() h.Stats.Close()
@ -222,8 +271,7 @@ func (h *Holder) Close() error {
} }
} }
if h.snapshotQueue != nil { if h.snapshotQueue != nil {
close(h.snapshotQueue) h.snapshotQueue.Stop()
// assuming the snapshotQueueWorker has already started, this is safe.
h.snapshotQueue = nil h.snapshotQueue = nil
} }

View file

@ -26,6 +26,7 @@ import (
"github.com/pilosa/pilosa/v2" "github.com/pilosa/pilosa/v2"
"github.com/pilosa/pilosa/v2/test" "github.com/pilosa/pilosa/v2/test"
"github.com/pkg/errors"
) )
func TestHolder_Open(t *testing.T) { func TestHolder_Open(t *testing.T) {
@ -197,7 +198,84 @@ func TestHolder_Open(t *testing.T) {
t.Fatalf("unexpected error: %s", err) t.Fatalf("unexpected error: %s", err)
} }
}) })
t.Run("ErrFragmentStorageRecoverable", func(t *testing.T) {
h := test.MustOpenHolder()
defer h.Close()
if idx, err := h.CreateIndex("foo", pilosa.IndexOptions{}); err != nil {
t.Fatal(err)
} else if field, err := idx.CreateField("bar", pilosa.OptFieldTypeDefault()); err != nil {
t.Fatal(err)
} else if _, err := field.SetBit(0, 0, nil); err != nil {
t.Fatal(err)
} else if err := h.Holder.Close(); err != nil {
t.Fatal(err)
} else if err := os.Truncate(filepath.Join(h.Path, "foo", "bar", "views", "standard", "fragments", "0"), 20); err != nil {
t.Fatal(err)
}
if err := h.Reopen(); err != nil {
t.Fatalf("unexpected error: %s", err)
}
})
t.Run("ForeignIndex", func(t *testing.T) {
t.Run("ErrForeignIndexNotFound", func(t *testing.T) {
h := test.MustOpenHolder()
defer h.Close()
if idx, err := h.CreateIndex("foo", pilosa.IndexOptions{}); err != nil {
t.Fatal(err)
} else {
_, err := idx.CreateField("bar", pilosa.OptFieldTypeInt(0, 100), pilosa.OptFieldForeignIndex("nonexistent"))
if err == nil {
t.Fatalf("expected error: %s", pilosa.ErrForeignIndexNotFound)
} else if errors.Cause(err) != pilosa.ErrForeignIndexNotFound {
t.Fatalf("expected error: %s, but got: %s", pilosa.ErrForeignIndexNotFound, err)
}
}
})
// Foreign index zzz is opened after foo/bar.
t.Run("ForeignIndexNotOpenYet", func(t *testing.T) {
h := test.MustOpenHolder()
defer h.Close()
if _, err := h.CreateIndex("zzz", pilosa.IndexOptions{}); err != nil {
t.Fatal(err)
} else if idx, err := h.CreateIndex("foo", pilosa.IndexOptions{}); err != nil {
t.Fatal(err)
} else if _, err := idx.CreateField("bar", pilosa.OptFieldTypeInt(0, 100), pilosa.OptFieldForeignIndex("zzz")); err != nil {
t.Fatal(err)
} else if err := h.Holder.Close(); err != nil {
t.Fatal(err)
}
if err := h.Reopen(); err != nil {
t.Fatalf("unexpected error: %s", err)
}
})
// Foreign index aaa is opened before foo/bar.
t.Run("ForeignIndexIsOpen", func(t *testing.T) {
h := test.MustOpenHolder()
defer h.Close()
if _, err := h.CreateIndex("aaa", pilosa.IndexOptions{}); err != nil {
t.Fatal(err)
} else if idx, err := h.CreateIndex("foo", pilosa.IndexOptions{}); err != nil {
t.Fatal(err)
} else if _, err := idx.CreateField("bar", pilosa.OptFieldTypeInt(0, 100), pilosa.OptFieldForeignIndex("aaa")); err != nil {
t.Fatal(err)
} else if err := h.Holder.Close(); err != nil {
t.Fatal(err)
}
if err := h.Reopen(); err != nil {
t.Fatalf("unexpected error: %s", err)
}
})
})
} }
func TestHolder_HasData(t *testing.T) { func TestHolder_HasData(t *testing.T) {

View file

@ -553,6 +553,35 @@ func (c *InternalClient) ImportValue(ctx context.Context, index, field string, s
return nil return nil
} }
// ImportValue2 is a simplified ImportValue method which just uses the
// ImportValueRequest instead of splitting up ImportValue and
// ImportValueK... it also supports importing float values. The idea
// being that (assuming it works) this will become the default (and be
// renamed) for 2.0, and we can deprecate the other methods.
func (c *InternalClient) ImportValue2(ctx context.Context, req *pilosa.ImportValueRequest, options *pilosa.ImportOptions) error {
span, ctx := tracing.StartSpanFromContext(ctx, "InternalClient.NewImportValue")
defer span.Finish()
buf, err := c.serializer.Marshal(req)
if err != nil {
return errors.Errorf("marshal import request: %s", err)
}
// Retrieve a list of nodes that own the shard.
nodes, err := c.FragmentNodes(ctx, req.Index, req.Shard)
if err != nil {
return errors.Errorf("shard nodes: %s", err)
}
// Import to each node.
for _, node := range nodes {
if err := c.importNode(ctx, node, req.Index, req.Field, buf, options); err != nil {
return errors.Errorf("import node: host=%s, err=%s", node.URI, err)
}
}
return nil
}
// ImportValueK bulk imports keyed field values to a host. // ImportValueK bulk imports keyed field values to a host.
func (c *InternalClient) ImportValueK(ctx context.Context, index, field string, vals []pilosa.FieldValue, opts ...pilosa.ImportOption) error { func (c *InternalClient) ImportValueK(ctx context.Context, index, field string, vals []pilosa.FieldValue, opts ...pilosa.ImportOption) error {
span, ctx := tracing.StartSpanFromContext(ctx, "InternalClient.ImportValueK") span, ctx := tracing.StartSpanFromContext(ctx, "InternalClient.ImportValueK")
@ -667,6 +696,56 @@ func (c *InternalClient) ImportRoaring(ctx context.Context, uri *pilosa.URI, ind
return nil return nil
} }
// ImportColumnAttrs does bulk import of column attrs
func (c *InternalClient) ImportColumnAttrs(ctx context.Context, uri *pilosa.URI, index string, req *pilosa.ImportColumnAttrsRequest) error {
span, ctx := tracing.StartSpanFromContext(ctx, "InternalClient.ImportRoaring")
defer span.Finish()
if index == "" {
return pilosa.ErrIndexRequired
}
if uri == nil {
uri = c.defaultURI
}
url := fmt.Sprintf("%s/index/%s/import-column-attrs", uri, index)
// Marshal data to protobuf.
data, err := c.serializer.Marshal(req)
if err != nil {
return errors.Wrap(err, "marshal import-column-attrs request")
}
// Generate HTTP request.
httpReq, err := http.NewRequest("POST", url, bytes.NewBuffer(data))
if err != nil {
return errors.Wrap(err, "creating request")
}
httpReq.Header.Set("Content-Type", "application/x-protobuf")
httpReq.Header.Set("Accept", "application/x-protobuf")
httpReq.Header.Set("User-Agent", "pilosa/"+pilosa.Version)
// Execute request against the host.
resp, err := c.executeRequest(httpReq.WithContext(ctx))
if err != nil {
return err
}
defer resp.Body.Close()
dec := json.NewDecoder(resp.Body)
rbody := &pilosa.ImportResponse{}
err = dec.Decode(rbody)
// Decode can return EOF when no error occurred. helpful!
if err != nil && err != io.EOF {
return errors.Wrap(err, "decoding response body")
}
if rbody.Err != "" {
return errors.Wrap(errors.New(rbody.Err), "importing roaring")
}
return nil
}
// ExportCSV bulk exports data for a single shard from a host to CSV format. // ExportCSV bulk exports data for a single shard from a host to CSV format.
func (c *InternalClient) ExportCSV(ctx context.Context, index, field string, shard uint64, w io.Writer) error { func (c *InternalClient) ExportCSV(ctx context.Context, index, field string, shard uint64, w io.Writer) error {
span, ctx := tracing.StartSpanFromContext(ctx, "InternalClient.ExportCSV") span, ctx := tracing.StartSpanFromContext(ctx, "InternalClient.ExportCSV")
@ -791,18 +870,28 @@ func (c *InternalClient) CreateFieldWithOptions(ctx context.Context, index, fiel
} }
// convert pilosa.FieldOptions to fieldOptions // convert pilosa.FieldOptions to fieldOptions
//
// TODO this kind of sucks because it's one more place that needs
// changes when we change anything with field options (and there
// are a lot of places already). It's not clear to me that this is
// providing a lot of value, but I think this kind of validation
// should probably happen in the field anyway??
fieldOpt := fieldOptions{ fieldOpt := fieldOptions{
Type: opt.Type, Type: opt.Type,
Keys: &opt.Keys, Keys: &opt.Keys,
} }
if fieldOpt.Type == "set" { if fieldOpt.Type == pilosa.FieldTypeSet {
fieldOpt.CacheType = &opt.CacheType fieldOpt.CacheType = &opt.CacheType
fieldOpt.CacheSize = &opt.CacheSize fieldOpt.CacheSize = &opt.CacheSize
} else if fieldOpt.Type == "int" { } else if fieldOpt.Type == pilosa.FieldTypeInt {
fieldOpt.Min = &opt.Min fieldOpt.Min = &opt.Min
fieldOpt.Max = &opt.Max fieldOpt.Max = &opt.Max
} else if fieldOpt.Type == "time" { } else if fieldOpt.Type == pilosa.FieldTypeTime {
fieldOpt.TimeQuantum = &opt.TimeQuantum fieldOpt.TimeQuantum = &opt.TimeQuantum
} else if fieldOpt.Type == pilosa.FieldTypeDecimal {
fieldOpt.Min = &opt.Min
fieldOpt.Max = &opt.Max
fieldOpt.Scale = &opt.Scale
} }
// TODO: remove buf completely? (depends on whether importer needs to create specific field types) // TODO: remove buf completely? (depends on whether importer needs to create specific field types)

View file

@ -22,6 +22,7 @@ import (
"fmt" "fmt"
gohttp "net/http" gohttp "net/http"
"reflect" "reflect"
"strconv"
"testing" "testing"
"time" "time"
@ -147,7 +148,8 @@ func TestClient_MultiNode(t *testing.T) {
} }
// Test must return exactly N results. // Test must return exactly N results.
if len(result.Results[0].([]pilosa.Pair)) != topN { pairsField := result.Results[0].(*pilosa.PairsField)
if len(pairsField.Pairs) != topN {
t.Fatalf("unexpected number of TopN results: %s", spew.Sdump(result)) t.Fatalf("unexpected number of TopN results: %s", spew.Sdump(result))
} }
p := []pilosa.Pair{ p := []pilosa.Pair{
@ -157,7 +159,7 @@ func TestClient_MultiNode(t *testing.T) {
{ID: 99, Count: 7}} {ID: 99, Count: 7}}
// Valdidate the Top 4 result counts. // Valdidate the Top 4 result counts.
if !reflect.DeepEqual(result.Results[0].([]pilosa.Pair), p) { if !reflect.DeepEqual(pairsField.Pairs, p) {
t.Fatalf("Invalid TopN result set: %s", spew.Sdump(result)) t.Fatalf("Invalid TopN result set: %s", spew.Sdump(result))
} }
@ -394,6 +396,60 @@ func TestClient_Import(t *testing.T) {
} }
} }
// Ensure client can bulk import column attrs.
func TestClient_ImportColumnAttrs(t *testing.T) {
cluster := test.MustNewCluster(t, 2)
for _, c := range cluster {
c.Config.Cluster.ReplicaN = 2
}
err := cluster.Start()
if err != nil {
t.Fatalf("starting cluster: %v", err)
}
defer cluster.Close()
ctx := context.Background()
_, err = cluster[0].API.CreateIndex(ctx, "i", pilosa.IndexOptions{})
if err != nil {
t.Fatalf("creating index: %v", err)
}
_, err = cluster[0].API.CreateField(ctx, "i", "f", pilosa.OptFieldTypeSet(pilosa.CacheTypeRanked, 100))
if err != nil {
t.Fatalf("creating field: %v", err)
}
_, err = cluster[0].API.Query(ctx, &pilosa.QueryRequest{Index: "i", Query: "Set(0, f=0) Set(1, f=0) Set(2, f=0) Set(3, f=0) Set(4, f=0)"})
if err != nil {
t.Fatalf("querying: %v", err)
}
attrKey := "k"
// Send import request.
host := cluster[0].URL()
c := MustNewClient(host, http.GetHTTPClient(nil))
colAttrsReq := makeImportColumnAttrsRequest("i", 0, attrKey)
if err := c.ImportColumnAttrs(ctx, &cluster[1].API.Node().URI, "i", colAttrsReq); err != nil {
t.Fatal(err)
}
// Verify data.
pql := "Options(Row(f=0), columnAttrs=true)"
res, err := cluster[1].API.Query(ctx, &pilosa.QueryRequest{Index: "i", Query: pql})
if err != nil {
t.Fatal(err)
}
if len(res.ColumnAttrSets) != 5 {
t.Fatal("incorrect number of column attrs set")
}
for _, v := range res.ColumnAttrSets {
attrVal := attrFun(v.ID)
if attrVal != v.Attrs[attrKey] {
t.Fatal(err)
}
}
}
// Ensure client can bulk import data. // Ensure client can bulk import data.
func TestClient_ImportRoaring(t *testing.T) { func TestClient_ImportRoaring(t *testing.T) {
cluster := test.MustNewCluster(t, 2) cluster := test.MustNewCluster(t, 2)
@ -550,14 +606,14 @@ func TestClient_ImportKeys(t *testing.T) {
Index: "keyed", Index: "keyed",
Query: "TopN(keyedf)", Query: "TopN(keyedf)",
}) })
if pairs, ok := resp.Results[0].([]pilosa.Pair); !ok { if pairs, ok := resp.Results[0].(*pilosa.PairsField); !ok {
t.Fatalf("unexpected response type %T", resp.Results[0]) t.Fatalf("unexpected response type %T", resp.Results[0])
} else if !reflect.DeepEqual(pairs, []pilosa.Pair{ } else if !reflect.DeepEqual(pairs.Pairs, []pilosa.Pair{
{Key: "green", Count: 3}, {Key: "green", Count: 3},
{Key: "blue", Count: 2}, {Key: "blue", Count: 2},
{Key: "purple", Count: 1}, {Key: "purple", Count: 1},
}) { }) {
t.Fatalf("unexpected topn result: %v", pairs) t.Fatalf("unexpected topn result: %v", pairs.Pairs)
} }
}) })
@ -577,14 +633,14 @@ func TestClient_ImportKeys(t *testing.T) {
Index: "keyed", Index: "keyed",
Query: "TopN(unkeyedf)", Query: "TopN(unkeyedf)",
}) })
if pairs, ok := resp.Results[0].([]pilosa.Pair); !ok { if pairs, ok := resp.Results[0].(*pilosa.PairsField); !ok {
t.Fatalf("unexpected response type %T", resp.Results[0]) t.Fatalf("unexpected response type %T", resp.Results[0])
} else if !reflect.DeepEqual(pairs, []pilosa.Pair{ } else if !reflect.DeepEqual(pairs.Pairs, []pilosa.Pair{
{ID: 1, Count: 3}, {ID: 1, Count: 3},
{ID: 2, Count: 2}, {ID: 2, Count: 2},
{ID: 3, Count: 1}, {ID: 3, Count: 1},
}) { }) {
t.Fatalf("unexpected topn result: %v", pairs) t.Fatalf("unexpected topn result: %v", pairs.Pairs)
} }
}) })
@ -604,14 +660,14 @@ func TestClient_ImportKeys(t *testing.T) {
Index: "unkeyed", Index: "unkeyed",
Query: "TopN(keyedf)", Query: "TopN(keyedf)",
}) })
if pairs, ok := resp.Results[0].([]pilosa.Pair); !ok { if pairs, ok := resp.Results[0].(*pilosa.PairsField); !ok {
t.Fatalf("unexpected response type %T", resp.Results[0]) t.Fatalf("unexpected response type %T", resp.Results[0])
} else if !reflect.DeepEqual(pairs, []pilosa.Pair{ } else if !reflect.DeepEqual(pairs.Pairs, []pilosa.Pair{
{Key: "green", Count: 3}, {Key: "green", Count: 3},
{Key: "blue", Count: 2}, {Key: "blue", Count: 2},
{Key: "purple", Count: 1}, {Key: "purple", Count: 1},
}) { }) {
t.Fatalf("unexpected topn result: %v", pairs) t.Fatalf("unexpected topn result: %v", pairs.Pairs)
} }
}) })
}) })
@ -649,14 +705,14 @@ func TestClient_ImportKeys(t *testing.T) {
Index: "keyed", Index: "keyed",
Query: "TopN(keyedf0)", Query: "TopN(keyedf0)",
}) })
if pairs, ok := resp.Results[0].([]pilosa.Pair); !ok { if pairs, ok := resp.Results[0].(*pilosa.PairsField); !ok {
t.Fatalf("unexpected response type %T", resp.Results[0]) t.Fatalf("unexpected response type %T", resp.Results[0])
} else if !reflect.DeepEqual(pairs, []pilosa.Pair{ } else if !reflect.DeepEqual(pairs.Pairs, []pilosa.Pair{
{Key: "green", Count: 3}, {Key: "green", Count: 3},
{Key: "blue", Count: 2}, {Key: "blue", Count: 2},
{Key: "purple", Count: 1}, {Key: "purple", Count: 1},
}) { }) {
t.Fatalf("unexpected topn result: %v", pairs) t.Fatalf("unexpected topn result: %v", pairs.Pairs)
} }
}) })
@ -681,14 +737,14 @@ func TestClient_ImportKeys(t *testing.T) {
Index: "keyed", Index: "keyed",
Query: "TopN(keyedf1)", Query: "TopN(keyedf1)",
}) })
if pairs, ok := resp.Results[0].([]pilosa.Pair); !ok { if pairs, ok := resp.Results[0].(*pilosa.PairsField); !ok {
t.Fatalf("unexpected response type %T", resp.Results[0]) t.Fatalf("unexpected response type %T", resp.Results[0])
} else if !reflect.DeepEqual(pairs, []pilosa.Pair{ } else if !reflect.DeepEqual(pairs.Pairs, []pilosa.Pair{
{Key: "green", Count: 3}, {Key: "green", Count: 3},
{Key: "blue", Count: 2}, {Key: "blue", Count: 2},
{Key: "purple", Count: 1}, {Key: "purple", Count: 1},
}) { }) {
t.Fatalf("unexpected topn result: %#v", pairs) t.Fatalf("unexpected topn result: %#v", pairs.Pairs)
} }
}) })
}) })
@ -777,6 +833,72 @@ func TestClient_ImportKeys(t *testing.T) {
}) })
} }
func TestClient_ImportIDs(t *testing.T) {
// Ensure that running a query between two imports does
// not affect the result set. It turns out, this is caused
// by the fragment.rowCache failing to be cleared after an
// importValue. This ensures that the rowCache is cleared
// after an import.
t.Run("ImportRangeImport", func(t *testing.T) {
cluster := test.MustRunCluster(t, 1)
defer cluster.Close()
cmd := cluster[0]
host := cmd.URL()
holder := cmd.Server.Holder()
hldr := test.Holder{Holder: holder}
idxName := "i"
fldName := "f"
// Load bitmap into cache to ensure cache gets updated.
index := hldr.MustCreateIndexIfNotExists(idxName, pilosa.IndexOptions{Keys: false})
_, err := index.CreateFieldIfNotExists(fldName, pilosa.OptFieldTypeInt(-10000, 10000))
if err != nil {
t.Fatal(err)
}
// Send import request.
c := MustNewClient(host, http.GetHTTPClient(nil))
if err := c.ImportValue(context.Background(), idxName, fldName, 0, []pilosa.FieldValue{
{ColumnID: 2, Value: 1},
}); err != nil {
t.Fatal(err)
}
// Verify range.
queryRequest := &pilosa.QueryRequest{
Query: fmt.Sprintf(`Row(%s>0)`, fldName),
Remote: false,
}
if result, err := c.Query(context.Background(), idxName, queryRequest); err != nil {
t.Fatal(err)
} else {
res := result.Results[0].(*pilosa.Row).Columns()
if !reflect.DeepEqual(res, []uint64{2}) {
t.Fatalf("unexpected column ids: %v", res)
}
}
// Send import request.
if err := c.ImportValue(context.Background(), idxName, fldName, 0, []pilosa.FieldValue{
{ColumnID: 1000, Value: 1},
}); err != nil {
t.Fatal(err)
}
// Verify range.
if result, err := c.Query(context.Background(), idxName, queryRequest); err != nil {
t.Fatal(err)
} else {
res := result.Results[0].(*pilosa.Row).Columns()
if !reflect.DeepEqual(res, []uint64{2, 1000}) {
t.Fatalf("unexpected column ids: %v", res)
}
}
})
}
// Ensure client can bulk import value data. // Ensure client can bulk import value data.
func TestClient_ImportValue(t *testing.T) { func TestClient_ImportValue(t *testing.T) {
cluster := test.MustRunCluster(t, 1) cluster := test.MustRunCluster(t, 1)
@ -998,6 +1120,114 @@ func TestClient_FragmentBlocks(t *testing.T) {
} }
} }
func TestClient_CreateDecimalField(t *testing.T) {
cluster := test.MustRunCluster(t, 1)
defer cluster.Close()
cmd := cluster[0]
c := MustNewClient(cmd.URL(), http.GetHTTPClient(nil))
index := "cdf"
err := c.CreateIndex(context.Background(), index, pilosa.IndexOptions{})
if err != nil {
t.Fatalf("creating index: %v", err)
}
field := "dfield"
err = c.CreateFieldWithOptions(context.Background(), index, field, pilosa.FieldOptions{Type: pilosa.FieldTypeDecimal, Scale: 1, Min: -1000, Max: 1000})
if err != nil {
t.Fatalf("creating field: %v", err)
}
fld, err := cmd.API.Field(context.Background(), index, field)
if err != nil {
t.Fatalf("getting field: %v", err)
}
if fld.Options().Scale != 1 {
t.Fatalf("expected Scale 1, got: %+v", fld.Options())
}
err = c.ImportValue2(context.Background(), &pilosa.ImportValueRequest{Index: index, Field: field, ColumnIDs: []uint64{1, 2, 3}, Shard: 0, FloatValues: []float64{1.1, 2.2, 3.3}}, &pilosa.ImportOptions{})
if err != nil {
t.Fatalf("importing float values: %v", err)
}
// Integer predicate.
resp, err := c.Query(context.Background(), index, &pilosa.QueryRequest{Index: index, Query: "Row(dfield>2)"})
if err != nil {
t.Fatalf("querying: %v", err)
}
if !reflect.DeepEqual(resp.Results[0].(*pilosa.Row).Columns(), []uint64{2, 3}) {
t.Fatalf("unexpected results: %v", resp.Results[0].(*pilosa.Row).Columns())
}
// Float predicate.
resp, err = c.Query(context.Background(), index, &pilosa.QueryRequest{Index: index, Query: "Row(dfield>2.1)"})
if err != nil {
t.Fatalf("querying: %v", err)
}
if !reflect.DeepEqual(resp.Results[0].(*pilosa.Row).Columns(), []uint64{2, 3}) {
t.Fatalf("unexpected results: %v", resp.Results[0].(*pilosa.Row).Columns())
}
// Integer predicates.
resp, err = c.Query(context.Background(), index, &pilosa.QueryRequest{Index: index, Query: "Row(1<dfield<3)"})
if err != nil {
t.Fatalf("querying: %v", err)
}
if !reflect.DeepEqual(resp.Results[0].(*pilosa.Row).Columns(), []uint64{1, 2}) {
t.Fatalf("unexpected results: %v", resp.Results[0].(*pilosa.Row).Columns())
}
// Float predicates.
resp, err = c.Query(context.Background(), index, &pilosa.QueryRequest{Index: index, Query: "Row(1.1<dfield<3.3)"})
if err != nil {
t.Fatalf("querying: %v", err)
}
if !reflect.DeepEqual(resp.Results[0].(*pilosa.Row).Columns(), []uint64{2}) {
t.Fatalf("unexpected results: %v", resp.Results[0].(*pilosa.Row).Columns())
}
resp, err = c.Query(context.Background(), index, &pilosa.QueryRequest{Index: index, Query: "Row(1.1<=dfield<3.3)"})
if err != nil {
t.Fatalf("querying: %v", err)
}
if !reflect.DeepEqual(resp.Results[0].(*pilosa.Row).Columns(), []uint64{1, 2}) {
t.Fatalf("unexpected results: %v", resp.Results[0].(*pilosa.Row).Columns())
}
resp, err = c.Query(context.Background(), index, &pilosa.QueryRequest{Index: index, Query: "Row(1.1<dfield<=3.3)"})
if err != nil {
t.Fatalf("querying: %v", err)
}
if !reflect.DeepEqual(resp.Results[0].(*pilosa.Row).Columns(), []uint64{2, 3}) {
t.Fatalf("unexpected results: %v", resp.Results[0].(*pilosa.Row).Columns())
}
resp, err = c.Query(context.Background(), index, &pilosa.QueryRequest{Index: index, Query: "Row(dfield<3.3)"})
if err != nil {
t.Fatalf("querying: %v", err)
}
if !reflect.DeepEqual(resp.Results[0].(*pilosa.Row).Columns(), []uint64{1, 2}) {
t.Fatalf("unexpected results: %v", resp.Results[0].(*pilosa.Row).Columns())
}
resp, err = c.Query(context.Background(), index, &pilosa.QueryRequest{Index: index, Query: "Row(dfield>2.2)"})
if err != nil {
t.Fatalf("querying: %v", err)
}
if !reflect.DeepEqual(resp.Results[0].(*pilosa.Row).Columns(), []uint64{3}) {
t.Fatalf("unexpected results: %v", resp.Results[0].(*pilosa.Row).Columns())
}
resp, err = c.Query(context.Background(), index, &pilosa.QueryRequest{Index: index, Query: "Row(dfield>=2.2)"})
if err != nil {
t.Fatalf("querying: %v", err)
}
if !reflect.DeepEqual(resp.Results[0].(*pilosa.Row).Columns(), []uint64{2, 3}) {
t.Fatalf("unexpected results: %v", resp.Results[0].(*pilosa.Row).Columns())
}
}
// Client represents a test wrapper for pilosa.Client. // Client represents a test wrapper for pilosa.Client.
type Client struct { type Client struct {
*http.InternalClient *http.InternalClient
@ -1021,3 +1251,23 @@ func makeImportRoaringRequest(clear bool, viewData string) *pilosa.ImportRoaring
}, },
} }
} }
func attrFun(id uint64) string {
return strconv.FormatInt(int64(id), 10)
}
func makeImportColumnAttrsRequest(index string, shard int64, attrKey string) *pilosa.ImportColumnAttrsRequest {
colIDs := make([]uint64, 0, 5)
attrVals := make([]string, 0, 5)
for n := uint64(0); n < 5; n++ {
colIDs = append(colIDs, n)
attrVals = append(attrVals, attrFun(n))
}
return &pilosa.ImportColumnAttrsRequest{
Index: index,
Shard: shard,
AttrKey: attrKey,
ColumnIDs: colIDs,
AttrVals: attrVals,
}
}

View file

@ -187,7 +187,7 @@ func (h *Handler) populateValidators() {
h.validators["DeleteField"] = queryValidationSpecRequired() h.validators["DeleteField"] = queryValidationSpecRequired()
h.validators["PostImport"] = queryValidationSpecRequired().Optional("clear", "ignoreKeyCheck") h.validators["PostImport"] = queryValidationSpecRequired().Optional("clear", "ignoreKeyCheck")
h.validators["PostImportRoaring"] = queryValidationSpecRequired().Optional("remote", "clear") h.validators["PostImportRoaring"] = queryValidationSpecRequired().Optional("remote", "clear")
h.validators["PostQuery"] = queryValidationSpecRequired().Optional("shards", "columnAttrs", "excludeRowAttrs", "excludeColumns") h.validators["PostQuery"] = queryValidationSpecRequired().Optional("shards", "columnAttrs", "excludeRowAttrs", "excludeColumns", "profile")
h.validators["GetInfo"] = queryValidationSpecRequired() h.validators["GetInfo"] = queryValidationSpecRequired()
h.validators["RecalculateCaches"] = queryValidationSpecRequired() h.validators["RecalculateCaches"] = queryValidationSpecRequired()
h.validators["GetSchema"] = queryValidationSpecRequired() h.validators["GetSchema"] = queryValidationSpecRequired()
@ -286,6 +286,7 @@ func newRouter(handler *Handler) *mux.Router {
router.HandleFunc("/index/{index}", handler.handlePostIndex).Methods("POST").Name("PostIndex") router.HandleFunc("/index/{index}", handler.handlePostIndex).Methods("POST").Name("PostIndex")
router.HandleFunc("/index/{index}", handler.handleDeleteIndex).Methods("DELETE").Name("DeleteIndex") router.HandleFunc("/index/{index}", handler.handleDeleteIndex).Methods("DELETE").Name("DeleteIndex")
//router.HandleFunc("/index/{index}/field", handler.handleGetFields).Methods("GET") // Not implemented. //router.HandleFunc("/index/{index}/field", handler.handleGetFields).Methods("GET") // Not implemented.
router.HandleFunc("/index/{index}/import-column-attrs", handler.handlePostImportColumnAttrs).Methods("POST").Name("PostImportColumnAttrs")
router.HandleFunc("/index/{index}/field/{field}", handler.handlePostField).Methods("POST").Name("PostField") router.HandleFunc("/index/{index}/field/{field}", handler.handlePostField).Methods("POST").Name("PostField")
router.HandleFunc("/index/{index}/field/{field}", handler.handleDeleteField).Methods("DELETE").Name("DeleteField") router.HandleFunc("/index/{index}/field/{field}", handler.handleDeleteField).Methods("DELETE").Name("DeleteField")
router.HandleFunc("/index/{index}/field/{field}/import", handler.handlePostImport).Methods("POST").Name("PostImport") router.HandleFunc("/index/{index}/field/{field}/import", handler.handlePostImport).Methods("POST").Name("PostImport")
@ -776,7 +777,7 @@ func (h *Handler) handlePostField(w http.ResponseWriter, r *http.Request) {
switch req.Options.Type { switch req.Options.Type {
case pilosa.FieldTypeSet: case pilosa.FieldTypeSet:
fos = append(fos, pilosa.OptFieldTypeSet(*req.Options.CacheType, *req.Options.CacheSize)) fos = append(fos, pilosa.OptFieldTypeSet(*req.Options.CacheType, *req.Options.CacheSize))
case pilosa.FieldTypeInt: case pilosa.FieldTypeInt, pilosa.FieldTypeDecimal:
if req.Options.Min == nil { if req.Options.Min == nil {
min := int64(math.MinInt64) min := int64(math.MinInt64)
req.Options.Min = &min req.Options.Min = &min
@ -785,7 +786,15 @@ func (h *Handler) handlePostField(w http.ResponseWriter, r *http.Request) {
max := int64(math.MaxInt64) max := int64(math.MaxInt64)
req.Options.Max = &max req.Options.Max = &max
} }
if req.Options.Type == pilosa.FieldTypeDecimal {
scale := int64(0)
if req.Options.Scale != nil {
scale = *req.Options.Scale
}
fos = append(fos, pilosa.OptFieldTypeDecimal(scale, *req.Options.Min, *req.Options.Max))
} else {
fos = append(fos, pilosa.OptFieldTypeInt(*req.Options.Min, *req.Options.Max)) fos = append(fos, pilosa.OptFieldTypeInt(*req.Options.Min, *req.Options.Max))
}
case pilosa.FieldTypeTime: case pilosa.FieldTypeTime:
fos = append(fos, pilosa.OptFieldTypeTime(*req.Options.TimeQuantum, req.Options.NoStandardView)) fos = append(fos, pilosa.OptFieldTypeTime(*req.Options.TimeQuantum, req.Options.NoStandardView))
case pilosa.FieldTypeMutex: case pilosa.FieldTypeMutex:
@ -798,6 +807,9 @@ func (h *Handler) handlePostField(w http.ResponseWriter, r *http.Request) {
fos = append(fos, pilosa.OptFieldKeys()) fos = append(fos, pilosa.OptFieldKeys())
} }
} }
if req.Options.ForeignIndex != nil {
fos = append(fos, pilosa.OptFieldForeignIndex(*req.Options.ForeignIndex))
}
_, err = h.api.CreateField(r.Context(), indexName, fieldName, fos...) _, err = h.api.CreateField(r.Context(), indexName, fieldName, fos...)
if _, ok := err.(pilosa.BadRequestError); ok { if _, ok := err.(pilosa.BadRequestError); ok {
@ -819,9 +831,11 @@ type fieldOptions struct {
CacheSize *uint32 `json:"cacheSize,omitempty"` CacheSize *uint32 `json:"cacheSize,omitempty"`
Min *int64 `json:"min,omitempty"` Min *int64 `json:"min,omitempty"`
Max *int64 `json:"max,omitempty"` Max *int64 `json:"max,omitempty"`
Scale *int64 `json:"scale,omitempty"`
TimeQuantum *pilosa.TimeQuantum `json:"timeQuantum,omitempty"` TimeQuantum *pilosa.TimeQuantum `json:"timeQuantum,omitempty"`
Keys *bool `json:"keys,omitempty"` Keys *bool `json:"keys,omitempty"`
NoStandardView bool `json:"noStandardView,omitempty"` NoStandardView bool `json:"noStandardView,omitempty"`
ForeignIndex *string `json:"foreignIndex,omitempty"`
} }
func (o *fieldOptions) validate() error { func (o *fieldOptions) validate() error {
@ -849,14 +863,18 @@ func (o *fieldOptions) validate() error {
return pilosa.NewBadRequestError(errors.New("max does not apply to field type set")) return pilosa.NewBadRequestError(errors.New("max does not apply to field type set"))
} else if o.TimeQuantum != nil { } else if o.TimeQuantum != nil {
return pilosa.NewBadRequestError(errors.New("timeQuantum does not apply to field type set")) return pilosa.NewBadRequestError(errors.New("timeQuantum does not apply to field type set"))
} else if o.ForeignIndex != nil {
return pilosa.NewBadRequestError(errors.New("set field cannot be a foreign key"))
} }
case pilosa.FieldTypeInt: case pilosa.FieldTypeInt, pilosa.FieldTypeDecimal:
if o.CacheType != nil { if o.CacheType != nil {
return pilosa.NewBadRequestError(errors.New("cacheType does not apply to field type int")) return pilosa.NewBadRequestError(errors.New("cacheType does not apply to field type int"))
} else if o.CacheSize != nil { } else if o.CacheSize != nil {
return pilosa.NewBadRequestError(errors.New("cacheSize does not apply to field type int")) return pilosa.NewBadRequestError(errors.New("cacheSize does not apply to field type int"))
} else if o.TimeQuantum != nil { } else if o.TimeQuantum != nil {
return pilosa.NewBadRequestError(errors.New("timeQuantum does not apply to field type int")) return pilosa.NewBadRequestError(errors.New("timeQuantum does not apply to field type int"))
} else if o.ForeignIndex != nil && o.Type == pilosa.FieldTypeDecimal {
return pilosa.NewBadRequestError(errors.New("decimal field cannot be a foreign key"))
} }
case pilosa.FieldTypeTime: case pilosa.FieldTypeTime:
if o.CacheType != nil { if o.CacheType != nil {
@ -869,6 +887,8 @@ func (o *fieldOptions) validate() error {
return pilosa.NewBadRequestError(errors.New("max does not apply to field type time")) return pilosa.NewBadRequestError(errors.New("max does not apply to field type time"))
} else if o.TimeQuantum == nil { } else if o.TimeQuantum == nil {
return pilosa.NewBadRequestError(errors.New("timeQuantum is required for field type time")) return pilosa.NewBadRequestError(errors.New("timeQuantum is required for field type time"))
} else if o.ForeignIndex != nil {
return pilosa.NewBadRequestError(errors.New("time field cannot be a foreign key"))
} }
case pilosa.FieldTypeMutex: case pilosa.FieldTypeMutex:
if o.CacheType == nil { if o.CacheType == nil {
@ -883,6 +903,8 @@ func (o *fieldOptions) validate() error {
return pilosa.NewBadRequestError(errors.New("max does not apply to field type mutex")) return pilosa.NewBadRequestError(errors.New("max does not apply to field type mutex"))
} else if o.TimeQuantum != nil { } else if o.TimeQuantum != nil {
return pilosa.NewBadRequestError(errors.New("timeQuantum does not apply to field type mutex")) return pilosa.NewBadRequestError(errors.New("timeQuantum does not apply to field type mutex"))
} else if o.ForeignIndex != nil {
return pilosa.NewBadRequestError(errors.New("mutex field cannot be a foreign key"))
} }
case pilosa.FieldTypeBool: case pilosa.FieldTypeBool:
if o.CacheType != nil { if o.CacheType != nil {
@ -897,6 +919,8 @@ func (o *fieldOptions) validate() error {
return pilosa.NewBadRequestError(errors.New("timeQuantum does not apply to field type bool")) return pilosa.NewBadRequestError(errors.New("timeQuantum does not apply to field type bool"))
} else if o.Keys != nil { } else if o.Keys != nil {
return pilosa.NewBadRequestError(errors.New("keys does not apply to field type bool")) return pilosa.NewBadRequestError(errors.New("keys does not apply to field type bool"))
} else if o.ForeignIndex != nil {
return pilosa.NewBadRequestError(errors.New("bool field cannot be a foreign key"))
} }
default: default:
return errors.Errorf("invalid field type: %s", o.Type) return errors.Errorf("invalid field type: %s", o.Type)
@ -1021,9 +1045,20 @@ func (h *Handler) readURLQueryRequest(r *http.Request) (*pilosa.QueryRequest, er
return nil, errors.New("invalid shard argument") return nil, errors.New("invalid shard argument")
} }
// Optional profiling
profile := false
profileString := q.Get("profile")
if profileString != "" {
profile, err = strconv.ParseBool(q.Get("profile"))
if err != nil {
return nil, fmt.Errorf("invalid profile argument: '%s' (should be true/false)", profileString)
}
}
return &pilosa.QueryRequest{ return &pilosa.QueryRequest{
Query: query, Query: query,
Shards: shards, Shards: shards,
Profile: profile,
ColumnAttrs: q.Get("columnAttrs") == "true", ColumnAttrs: q.Get("columnAttrs") == "true",
ExcludeRowAttrs: q.Get("excludeRowAttrs") == "true", ExcludeRowAttrs: q.Get("excludeRowAttrs") == "true",
ExcludeColumns: q.Get("excludeColumns") == "true", ExcludeColumns: q.Get("excludeColumns") == "true",
@ -1101,7 +1136,7 @@ func (h *Handler) handlePostImport(w http.ResponseWriter, r *http.Request) {
} }
// Unmarshal request based on field type. // Unmarshal request based on field type.
if field.Type() == pilosa.FieldTypeInt { if field.Type() == pilosa.FieldTypeInt || field.Type() == pilosa.FieldTypeDecimal {
// Field type: Int // Field type: Int
// Marshal into request object. // Marshal into request object.
req := &pilosa.ImportValueRequest{} req := &pilosa.ImportValueRequest{}
@ -1601,6 +1636,50 @@ func GetHTTPClient(t *tls.Config) *http.Client {
return &http.Client{Transport: transport} return &http.Client{Transport: transport}
} }
// handlePostImportColumnAttrs
func (h *Handler) handlePostImportColumnAttrs(w http.ResponseWriter, r *http.Request) {
// Verify that request is only communicating over protobufs.
if r.Header.Get("Content-Type") != "application/x-protobuf" {
http.Error(w, "Unsupported media type", http.StatusUnsupportedMediaType)
return
} else if r.Header.Get("Accept") != "application/x-protobuf" {
http.Error(w, "Not acceptable", http.StatusNotAcceptable)
return
}
opts := []pilosa.ImportOption{}
body, err := ioutil.ReadAll(r.Body)
if err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
req := &pilosa.ImportColumnAttrsRequest{}
if err := h.api.Serializer.Unmarshal(body, req); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
if err := h.api.ImportColumnAttrs(r.Context(), req, opts...); err != nil {
http.Error(w, err.Error(), http.StatusInternalServerError)
return
}
// Marshal response object.
buf, e := h.api.Serializer.Marshal(&pilosa.ImportResponse{Err: ""})
if e != nil {
http.Error(w, fmt.Sprintf("marshal import-column-attrs response"), http.StatusInternalServerError)
return
}
// Write response.
_, err = w.Write(buf)
if err != nil {
h.logger.Printf("writing import-column-attrs response: %v", err)
}
}
// handlPostRoaringImport // handlPostRoaringImport
func (h *Handler) handlePostImportRoaring(w http.ResponseWriter, r *http.Request) { func (h *Handler) handlePostImportRoaring(w http.ResponseWriter, r *http.Request) {
// Verify that request is only communicating over protobufs. // Verify that request is only communicating over protobufs.
@ -1665,7 +1744,7 @@ func (h *Handler) handlePostImportRoaring(w http.ResponseWriter, r *http.Request
// Marshal response object. // Marshal response object.
buf, err := h.api.Serializer.Marshal(resp) buf, err := h.api.Serializer.Marshal(resp)
if err != nil { if err != nil {
http.Error(w, fmt.Sprintf("marshal import response: %v", err), http.StatusInternalServerError) http.Error(w, fmt.Sprintf("marshal import-roaring response: %v", err), http.StatusInternalServerError)
return return
} }

View file

@ -58,9 +58,10 @@ type Index struct {
Stats stats.StatsClient Stats stats.StatsClient
logger logger.Logger logger logger.Logger
snapshotQueue chan *fragment snapshotQueue snapshotQueue
// Used for notifying holder when a field is added. // Used for notifying holder when a field is added.
// Also passed to field for foreign-index lookup.
holder *Holder holder *Holder
// Instantiates new translation stores for fields. // Instantiates new translation stores for fields.
@ -197,6 +198,10 @@ fileLoop:
return errors.Wrapf(ErrName, "'%s'", fi.Name()) return errors.Wrapf(ErrName, "'%s'", fi.Name())
} }
// Pass holder through to the field for use in looking
// up a foreign index.
fld.holder = i.holder
if err := fld.Open(); err != nil { if err := fld.Open(); err != nil {
return fmt.Errorf("open field: name=%s, err=%s", fld.Name(), err) return fmt.Errorf("open field: name=%s, err=%s", fld.Name(), err)
} }
@ -426,17 +431,17 @@ func (i *Index) createField(name string, opt FieldOptions) (*Field, error) {
return nil, errors.Wrap(err, "initializing") return nil, errors.Wrap(err, "initializing")
} }
// Pass holder through to the field for use in looking
// up a foreign index.
f.holder = i.holder
f.setOptions(&opt)
// Open field. // Open field.
if err := f.Open(); err != nil { if err := f.Open(); err != nil {
return nil, errors.Wrap(err, "opening") return nil, errors.Wrap(err, "opening")
} }
// Apply field options.
if err := f.applyOptions(opt); err != nil {
f.Close()
return nil, errors.Wrap(err, "applying options")
}
if err := f.saveMeta(); err != nil { if err := f.saveMeta(); err != nil {
f.Close() f.Close()
return nil, errors.Wrap(err, "saving meta") return nil, errors.Wrap(err, "saving meta")
@ -462,7 +467,9 @@ func (i *Index) newField(path, name string) (*Field, error) {
f.Stats = i.Stats f.Stats = i.Stats
f.broadcaster = i.broadcaster f.broadcaster = i.broadcaster
f.rowAttrStore = i.newAttrStore(filepath.Join(f.path, ".data")) f.rowAttrStore = i.newAttrStore(filepath.Join(f.path, ".data"))
if i.snapshotQueue != nil {
f.snapshotQueue = i.snapshotQueue f.snapshotQueue = i.snapshotQueue
}
f.OpenTranslateStore = i.OpenTranslateStore f.OpenTranslateStore = i.OpenTranslateStore
return f, nil return f, nil
} }

View file

@ -51,6 +51,6 @@ services:
volumes: volumes:
- /var/run/docker.sock:/var/run/docker.sock - /var/run/docker.sock:/var/run/docker.sock
command: command:
- "cd /go/src/github.com/pilosa/pilosa/ && go test -v -count=1 github.com/pilosa/pilosa/internal/clustertests" - "cd /go/src/github.com/pilosa/pilosa/ && go test -mod=vendor -v -count=1 github.com/pilosa/pilosa/v2/internal/clustertests"
networks: networks:
pilosanet: pilosanet:

File diff suppressed because it is too large Load diff

View file

@ -18,6 +18,8 @@ message FieldOptions {
bool NoStandardView = 12; bool NoStandardView = 12;
int64 Base = 13; int64 Base = 13;
uint64 BitDepth = 14; uint64 BitDepth = 14;
int64 Scale = 15;
string ForeignIndex = 16;
} }
message ImportResponse { message ImportResponse {

File diff suppressed because it is too large Load diff

View file

@ -6,6 +6,12 @@ message Row {
repeated uint64 Columns = 1; repeated uint64 Columns = 1;
repeated string Keys = 3; repeated string Keys = 3;
repeated Attr Attrs = 2; repeated Attr Attrs = 2;
bytes Roaring = 4;
}
message SignedRow {
Row Pos = 1;
Row Neg = 2;
} }
message RowIdentifiers { message RowIdentifiers {
@ -19,6 +25,16 @@ message Pair {
uint64 Count = 2; uint64 Count = 2;
} }
message PairField {
Pair Pair = 1;
string Field = 2;
}
message PairsField {
repeated Pair Pairs = 1;
string Field = 2;
}
message FieldRow{ message FieldRow{
string Field = 1; string Field = 1;
uint64 RowID = 2; uint64 RowID = 2;
@ -28,6 +44,7 @@ message FieldRow{
message GroupCount{ message GroupCount{
repeated FieldRow Group = 1; repeated FieldRow Group = 1;
uint64 Count = 2; uint64 Count = 2;
int64 Sum = 3;
} }
message ValCount { message ValCount {
@ -61,6 +78,7 @@ message QueryRequest {
bool Remote = 5; bool Remote = 5;
bool ExcludeRowAttrs = 6; bool ExcludeRowAttrs = 6;
bool ExcludeColumns = 7; bool ExcludeColumns = 7;
repeated Row EmbeddedData = 8;
} }
message QueryResponse { message QueryResponse {
@ -79,6 +97,8 @@ message QueryResult {
repeated uint64 RowIDs = 7; repeated uint64 RowIDs = 7;
repeated GroupCount GroupCounts = 8; repeated GroupCount GroupCounts = 8;
RowIdentifiers RowIdentifiers = 9; RowIdentifiers RowIdentifiers = 9;
SignedRow SignedRow = 10;
PairsField PairsField = 11;
} }
message ImportRequest { message ImportRequest {
@ -99,6 +119,8 @@ message ImportValueRequest {
repeated uint64 ColumnIDs = 5; repeated uint64 ColumnIDs = 5;
repeated string ColumnKeys = 7; repeated string ColumnKeys = 7;
repeated int64 Values = 6; repeated int64 Values = 6;
repeated double FloatValues = 8;
repeated string StringValues = 9;
} }
message TranslateKeysRequest { message TranslateKeysRequest {
@ -120,3 +142,11 @@ message ImportRoaringRequest {
bool Clear = 1; bool Clear = 1;
repeated ImportRoaringRequestView views = 2; repeated ImportRoaringRequestView views = 2;
} }
message ImportColumnAttrsRequest {
string Index = 1;
int64 Shard = 2;
string AttrKey = 3;
repeated string AttrVals = 4;
repeated uint64 ColumnIDs = 5;
}

10
license.exceptions Normal file
View file

@ -0,0 +1,10 @@
# names of files which we do not expect to have our license header
./apimethod_string.go
./pql/pql.peg.go
./internal/private.pb.go
./internal/public.pb.go
./lru/lru.go
./enterprise/enterprise.go
./roaring/btree.go
./roaring/btree_test.go
./proto/pilosa.pb.go

View file

@ -29,6 +29,8 @@ var (
ErrIndexExists = errors.New("index already exists") ErrIndexExists = errors.New("index already exists")
ErrIndexNotFound = errors.New("index not found") ErrIndexNotFound = errors.New("index not found")
ErrForeignIndexNotFound = errors.New("foreign index not found")
// ErrFieldRequired is returned when no field is specified. // ErrFieldRequired is returned when no field is specified.
ErrFieldRequired = errors.New("field required") ErrFieldRequired = errors.New("field required")
ErrFieldExists = errors.New("field already exists") ErrFieldExists = errors.New("field already exists")
@ -48,7 +50,7 @@ var (
ErrInvalidView = errors.New("invalid view") ErrInvalidView = errors.New("invalid view")
ErrInvalidCacheType = errors.New("invalid cache type") ErrInvalidCacheType = errors.New("invalid cache type")
ErrName = errors.New("invalid index or field name, must match [a-z][a-z0-9_-]* and contain at most 64 characters") ErrName = errors.New("invalid index or field name, must match [a-z][a-z0-9_-]* and contain at most 230 characters")
ErrLabel = errors.New("invalid row or column label, must match [A-Za-z0-9_-]") ErrLabel = errors.New("invalid row or column label, must match [A-Za-z0-9_-]")
// ErrFragmentNotFound is returned when a fragment does not exist. // ErrFragmentNotFound is returned when a fragment does not exist.
@ -118,7 +120,7 @@ func newNotFoundError(err error) NotFoundError {
} }
// Regular expression to validate index and field names. // Regular expression to validate index and field names.
var nameRegexp = regexp.MustCompile(`^[a-z][a-z0-9_-]{0,63}$`) var nameRegexp = regexp.MustCompile(`^[a-z][a-z0-9_-]{0,229}$`)
// ColumnAttrSet represents a set of attributes for a vertical column in an index. // ColumnAttrSet represents a set of attributes for a vertical column in an index.
// Can have a set of attributes attached to it. // Can have a set of attributes attached to it.

View file

@ -21,7 +21,7 @@ import (
func TestValidateName(t *testing.T) { func TestValidateName(t *testing.T) {
names := []string{ names := []string{
"a", "ab", "ab1", "b-c", "d_e", "exists", "a", "ab", "ab1", "b-c", "d_e", "exists",
"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", "longbutnottoolongaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa12345689012345689012345678901234567890",
} }
for _, name := range names { for _, name := range names {
if validateName(name) != nil { if validateName(name) != nil {
@ -33,7 +33,7 @@ func TestValidateName(t *testing.T) {
func TestValidateNameInvalid(t *testing.T) { func TestValidateNameInvalid(t *testing.T) {
names := []string{ names := []string{
"", "'", "^", "/", "\\", "A", "*", "a:b", "valid?no", "yüce", "1", "_", "-", "", "'", "^", "/", "\\", "A", "*", "a:b", "valid?no", "yüce", "1", "_", "-",
"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa1", "_exists", "long123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa1", "_exists",
} }
for _, name := range names { for _, name := range names {
if validateName(name) == nil { if validateName(name) == nil {

View file

@ -17,10 +17,13 @@ package pql
import ( import (
"bytes" "bytes"
"fmt" "fmt"
"reflect"
"sort" "sort"
"strconv" "strconv"
"strings" "strings"
"time" "time"
"github.com/molecula/ext"
) )
// Query represents a PQL query. // Query represents a PQL query.
@ -84,19 +87,26 @@ func (q *Query) endConditional() {
if len(q.conditional) != 5 { if len(q.conditional) != 5 {
panic(fmt.Sprintf("conditional of wrong length: %#v", q.conditional)) panic(fmt.Sprintf("conditional of wrong length: %#v", q.conditional))
} }
low, _ := strconv.ParseInt(q.conditional[0], 10, 64) low := parseNum(q.conditional[0])
field := q.conditional[2] field := q.conditional[2]
high, _ := strconv.ParseInt(q.conditional[4], 10, 64) high := parseNum(q.conditional[4])
if q.conditional[1] == "<" { var op Token
low++ switch q.conditional[1] + q.conditional[3] {
} case "<<":
if q.conditional[3] == "<" { op = BTWN_LT_LT
high-- case "<=<":
op = BTWN_LTE_LT
case "<<=":
op = BTWN_LT_LTE
case "<=<=":
op = BETWEEN
default:
panic(fmt.Sprintf("impossible conditional ops: '%s' and '%s'", q.conditional[1], q.conditional[3]))
} }
elem := q.lastCallStackElem() elem := q.lastCallStackElem()
elem.call.Args[field] = &Condition{Op: BETWEEN, Value: []interface{}{low, high}} elem.call.Args[field] = &Condition{Op: op, Value: []interface{}{low, high}}
q.conditional = nil q.conditional = nil
} }
@ -152,16 +162,7 @@ func (q *Query) addNumVal(val string) {
if elem == nil || elem.lastField == "" { if elem == nil || elem.lastField == "" {
panic(fmt.Sprintf("addIntVal called with '%s' when lastField is empty", val)) panic(fmt.Sprintf("addIntVal called with '%s' when lastField is empty", val))
} }
var ival interface{} ival := parseNum(val)
var err error
if strings.Contains(val, ".") {
ival, err = strconv.ParseFloat(val, 64)
} else {
ival, err = strconv.ParseInt(val, 10, 64)
}
if err != nil {
panic(fmt.Sprintf("%s: %s", intOutOfRangeError, err))
}
if elem.inList { if elem.inList {
if elem.lastCond != ILLEGAL { if elem.lastCond != ILLEGAL {
list := elem.call.Args[elem.lastField].(*Condition).Value.([]interface{}) list := elem.call.Args[elem.lastField].(*Condition).Value.([]interface{})
@ -259,11 +260,259 @@ type callStackElem struct {
inList bool inList bool
} }
// Call represents a function call in the AST. // Some call types may require special handling, which needs to occur
// before distributing processing to individual shards.
type CallType byte
const (
// Normal calls can be executed per shard.
PrecallNone = CallType(iota)
// PreCallGlobal indicates a call which must be run globally *before*
// distributing the call to other shards. Example: A Distinct query,
// where every shard could potentially produce results for any shard,
// so you have to produce the results up front.
PrecallGlobal
// PreCallPerNode indicates a call which needs to be run per-shard
// in a way that lets it be done on each shard, but where it should
// be done prior to spawning per-shard goroutines. Example:
// A cross-index query, where each local shard may or may not need
// to get data from a remote node, but batches of shards can
// probably be gotten from the same remote node.
PrecallPerNode
)
// Call represents a function call in the AST. The Precomputed field
// is used by the executor to handle non-standard call types; it does
// these by actually executing them separately, then replacing them
// in the call tree with a new call using the special precomputed
// type, with the Precomputed field set to a map from shards to results.
type Call struct { type Call struct {
Name string Name string
Args map[string]interface{} Args map[string]interface{}
Children []*Call Children []*Call
Type CallType
Precomputed map[uint64]interface{}
}
// callInfo defines the arguments allowed for a particular PQL call, and
// possibly things about its semantics. If allowUnknown is true, unfamiliar
// non-reserved names are allowed on the assumption that they're field names.
// Otherwise, only those names explicitly listed are allowed. Reserved args
// (those with a leading underscore) are never allowed unless explicitly
// present.
//
// The prototypes map maps from argument names to a value. If the value is
// non-nil, the argument will be checked for type-matching. So, for instance,
// `x: 10` would indicate that x must be an int.
type callInfo struct {
allowUnknown bool
prototypes map[string]interface{}
callType CallType
}
// We want to be able to accept either a string or int64 for
// field names. Special-case type:
type stringOrInt64Type struct{}
var stringOrInt64 stringOrInt64Type
var allowUnderField = callInfo{
allowUnknown: true,
prototypes: map[string]interface{}{
"_field": "",
},
}
var allowField = callInfo{
allowUnknown: false,
prototypes: map[string]interface{}{
"field": "",
},
}
var callInfoByFunc = map[string]callInfo{
// the easy cases: things that take arbitrary inputs, because they're
// taking field=value cases
"Bitmap": {allowUnknown: true},
"Count": {allowUnknown: true},
"Row": {allowUnknown: true},
"Range": {allowUnknown: true},
// allow only "field=X" cases with string field names
"Max": allowField,
"Min": allowField,
"Sum": allowField,
// only take other calls, should never have "args"
"Difference": {allowUnknown: false},
"Intersect": {allowUnknown: false},
"Not": {allowUnknown: false},
"All": {
allowUnknown: false,
prototypes: map[string]interface{}{
"limit": int64(0),
"offset": int64(0),
},
},
"ClearRow": {allowUnknown: true},
"Store": {allowUnknown: true},
"MinRow": allowField,
"MaxRow": allowField,
"Rows": {
allowUnknown: false,
prototypes: map[string]interface{}{
"_field": "",
"field": "",
"limit": int64(0),
"column": nil,
"previous": nil,
"from": nil,
"to": nil,
},
},
"Shift": {allowUnknown: false,
prototypes: map[string]interface{}{
"n": int64(0),
},
},
"Union": {allowUnknown: false},
"Xor": {allowUnknown: false},
// things that take _field
"TopN": allowUnderField,
// special cases:
"Clear": {
allowUnknown: true,
prototypes: map[string]interface{}{
"_col": stringOrInt64,
},
},
"GroupBy": {
allowUnknown: false,
prototypes: map[string]interface{}{
"filter": nil,
"limit": int64(0),
"previous": nil,
"aggregate": nil,
"having": nil,
},
},
"Options": {
allowUnknown: false,
prototypes: map[string]interface{}{
"excludeRowAttrs": true,
"excludeColumns": true,
"columnAttrs": true,
"shards": nil,
},
},
"Set": {
allowUnknown: true,
prototypes: map[string]interface{}{
"_col": stringOrInt64,
"_timestamp": "",
},
},
"Precomputed": {
allowUnknown: true,
},
"SetBit": {
allowUnknown: true,
prototypes: map[string]interface{}{
"_col": stringOrInt64,
},
},
"SetRowAttrs": {
allowUnknown: true,
prototypes: map[string]interface{}{
"_field": "",
"_row": stringOrInt64,
},
},
"SetColumnAttrs": {
allowUnknown: true,
prototypes: map[string]interface{}{
"_field": "",
"_col": stringOrInt64,
},
},
"IncludesColumn": {
allowUnknown: false,
prototypes: map[string]interface{}{
"column": stringOrInt64,
},
},
}
// RegisterPluginFuncs adds arg validation for plugin funcs. Not very good
// arg validation.
func RegisterPluginFuncs(ops []ext.BitmapOp) {
for _, op := range ops {
// ignore overlap for now. This should change.
if _, ok := callInfoByFunc[op.Name]; ok {
continue
}
ci := callInfo{allowUnknown: true}
if len(op.Reserved) > 0 {
// mark these as valid/known reserved words
ci.prototypes = make(map[string]interface{})
for _, res := range op.Reserved {
ci.prototypes[res] = nil
}
}
t := op.Func.BitmapOpType()
if t.Precall == ext.OpPrecallGlobal {
ci.callType = PrecallGlobal
}
callInfoByFunc[op.Name] = ci
}
}
// CheckCallInfo tries to validate that arguments are correct and valid for the
// given call. It does not guarantee checking all possible errors; for instance,
// if an argument is a field name, CheckCallInfo can't validate that the field
// exists. It also updates with information like whether the call is expected
// to require precalling.
func (c *Call) CheckCallInfo() error {
valid, ok := callInfoByFunc[c.Name]
if !ok {
return fmt.Errorf("no arg validation for '%s'", c.Name)
}
c.Type = valid.callType
for k, v := range c.Args {
acceptable, ok := valid.prototypes[k]
if !ok && !valid.allowUnknown {
return fmt.Errorf("'%s': unknown arg '%s'", c.String(), k)
}
if !ok && strings.HasPrefix(k, "_") {
return fmt.Errorf("'%s': unknown reserved arg '%s'", c.String(), k)
}
if acceptable == nil {
continue
}
// if the types are identical, that's fine
if reflect.TypeOf(acceptable) == reflect.TypeOf(v) {
continue
}
if reflect.TypeOf(acceptable) == reflect.TypeOf(stringOrInt64) {
switch v.(type) {
case string, int64:
continue
default:
return fmt.Errorf("'%s': arg '%s' needed a string or integer value, got %T.",
c.String(), k, v)
}
}
return fmt.Errorf("'%s': arg '%s' wrong type (got %T, expected %T)",
c.String(), k, v, acceptable)
}
// call-specific checking
for _, child := range c.Children {
if err := child.CheckCallInfo(); err != nil {
return err
}
}
return nil
} }
// FieldArg determines which key-value pair contains the field and rowID, // FieldArg determines which key-value pair contains the field and rowID,
@ -283,13 +532,29 @@ func IsReservedArg(name string) bool {
return true return true
} }
switch name { switch name {
case "from", "to": case "from", "to", "index":
return true return true
default: default:
return false return false
} }
} }
// CallIndex handles guessing whether we've been asked to apply this to a
// different index. An empty string means "no".
func (c *Call) CallIndex() string {
if index, ok := c.Args["_index"]; ok {
if index, ok := index.(string); ok {
return index
}
}
if index, ok := c.Args["index"]; ok && index != "" {
if index, ok := index.(string); ok {
return index
}
}
return ""
}
// BoolArg is for reading the value at key from call.Args as a bool. If the // BoolArg is for reading the value at key from call.Args as a bool. If the
// key is not in Call.Args, the value of the returned bool will be false, and // key is not in Call.Args, the value of the returned bool will be false, and
// the error will be nil. The value is assumed to be a bool. An error is // the error will be nil. The value is assumed to be a bool. An error is
@ -489,9 +754,39 @@ func (cond *Condition) String() string {
return fmt.Sprintf("%s %s", cond.Op.String(), formatValue(cond.Value)) return fmt.Sprintf("%s %s", cond.Op.String(), formatValue(cond.Value))
} }
// StringWithSubj returns the string representation of the condition
// including the provided subject.
func (cond *Condition) StringWithSubj(subj string) string {
switch cond.Op {
case EQ, NEQ, LT, LTE, GT, GTE:
return fmt.Sprintf("%s%s", subj, cond.String())
case BETWEEN, BTWN_LT_LTE, BTWN_LTE_LT, BTWN_LT_LT:
val, ok := cond.Int64SliceValue() // TODO: this should depend on subj type (int64 vs. uint64)
if !ok || len(val) < 2 {
return ""
}
if cond.Op == BETWEEN {
return fmt.Sprintf("%d<=%s<=%d", val[0], subj, val[1])
} else if cond.Op == BTWN_LT_LTE {
return fmt.Sprintf("%d<%s<=%d", val[0], subj, val[1])
} else if cond.Op == BTWN_LTE_LT {
return fmt.Sprintf("%d<=%s<%d", val[0], subj, val[1])
} else if cond.Op == BTWN_LT_LT {
return fmt.Sprintf("%d<%s<%d", val[0], subj, val[1])
}
}
return ""
}
// IntSliceValue reads cond.Value as a slice of uint64. // IntSliceValue reads cond.Value as a slice of uint64.
// If the value is a slice of uint64 it will convert // If the value is a slice of uint64 it will convert
// it to []int64. Otherwise, if it is not a []int64 it will return an error. // it to []int64. Otherwise, if it is not a []int64 it will return an error.
//
// TODO(2.0) this is now only referenced in a test and should probably
// be removed. The functionality was replaced by getCondIntSlice in
// pilosa/executor.go which needed to check for floating point values
// and also have access to the Pilosa field to see if floating point
// values were valid and how they needed to be scaled.
func (cond *Condition) IntSliceValue() ([]int64, error) { func (cond *Condition) IntSliceValue() ([]int64, error) {
val := cond.Value val := cond.Value
@ -514,6 +809,79 @@ func (cond *Condition) IntSliceValue() ([]int64, error) {
} }
} }
func (cond *Condition) Uint64Value() (uint64, bool) {
val := cond.Value
switch tval := val.(type) {
case int64:
if tval >= 0 {
return uint64(tval), true
}
case uint64:
return tval, true
}
return 0, false
}
func (cond *Condition) Uint64SliceValue() ([]uint64, bool) {
val := cond.Value
switch tval := val.(type) {
case []interface{}:
ret := make([]uint64, len(tval))
for i, v := range tval {
switch tv := v.(type) {
case int64:
ret[i] = uint64(tv)
case uint64:
ret[i] = tv
default:
return nil, false
}
}
return ret, true
}
return nil, false
}
func (cond *Condition) Int64Value() (int64, bool) {
val := cond.Value
switch tval := val.(type) {
case int64:
return tval, true
case uint64:
// TODO: consider overflow?
return int64(tval), true
}
return 0, false
}
func (cond *Condition) Int64SliceValue() ([]int64, bool) {
val := cond.Value
switch tval := val.(type) {
case []interface{}:
ret := make([]int64, len(tval))
for i, v := range tval {
switch tv := v.(type) {
case int64:
ret[i] = tv
case uint64:
ret[i] = int64(tv)
default:
return nil, false
}
}
return ret, true
}
return nil, false
}
func formatValue(v interface{}) string { func formatValue(v interface{}) string {
switch v := v.(type) { switch v := v.(type) {
case string: case string:
@ -560,3 +928,17 @@ func joinUint64Slice(a []uint64) string {
} }
return "[" + strings.Join(other, ",") + "]" return "[" + strings.Join(other, ",") + "]"
} }
func parseNum(val string) interface{} {
var ival interface{}
var err error
if strings.Contains(val, ".") {
ival, err = strconv.ParseFloat(val, 64)
} else {
ival, err = strconv.ParseInt(val, 10, 64)
}
if err != nil {
panic(fmt.Sprintf("%s: %s", intOutOfRangeError, err))
}
return ival
}

View file

@ -82,6 +82,14 @@ func (p *parser) Parse() (*Query, error) {
panic(v) panic(v)
} }
} }
for _, call := range p.Query.Calls {
if call == nil {
return nil, fmt.Errorf("unexpected nil Call in query's call list")
}
if err := call.CheckCallInfo(); err != nil {
return nil, err
}
}
return &p.Query, nil return &p.Query, nil
} }

View file

@ -75,12 +75,12 @@ func TestParser_Parse(t *testing.T) {
// Parse with only arguments. // Parse with only arguments.
t.Run("ArgumentsOnly", func(t *testing.T) { t.Run("ArgumentsOnly", func(t *testing.T) {
q, err := pql.ParseString(`MyCall( key= value, foo='bar', age = 12 , bool0=true, bool1=false, x=null, escape="\" \\escape\n\\\\" )`) q, err := pql.ParseString(`Row( key= value, foo='bar', age = 12 , bool0=true, bool1=false, x=null, escape="\" \\escape\n\\\\" )`)
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(q.Calls[0], } else if !reflect.DeepEqual(q.Calls[0],
&pql.Call{ &pql.Call{
Name: "MyCall", Name: "Row",
Args: map[string]interface{}{ Args: map[string]interface{}{
"key": "value", "key": "value",
"foo": "bar", "foo": "bar",
@ -98,12 +98,12 @@ func TestParser_Parse(t *testing.T) {
// Parse with float arguments. // Parse with float arguments.
t.Run("WithFloatArgs", func(t *testing.T) { t.Run("WithFloatArgs", func(t *testing.T) {
q, err := pql.ParseString(`MyCall( key=12.25, foo= 13.167, bar=2., baz=0.9)`) q, err := pql.ParseString(`Row( key=12.25, foo= 13.167, bar=2., baz=0.9)`)
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(q.Calls[0], } else if !reflect.DeepEqual(q.Calls[0],
&pql.Call{ &pql.Call{
Name: "MyCall", Name: "Row",
Args: map[string]interface{}{ Args: map[string]interface{}{
"key": 12.25, "key": 12.25,
"foo": 13.167, "foo": 13.167,
@ -118,12 +118,12 @@ func TestParser_Parse(t *testing.T) {
// Parse with float arguments. // Parse with float arguments.
t.Run("WithNegativeArgs", func(t *testing.T) { t.Run("WithNegativeArgs", func(t *testing.T) {
q, err := pql.ParseString(`MyCall( key=-12.25, foo= -13)`) q, err := pql.ParseString(`Row( key=-12.25, foo= -13)`)
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(q.Calls[0], } else if !reflect.DeepEqual(q.Calls[0],
&pql.Call{ &pql.Call{
Name: "MyCall", Name: "Row",
Args: map[string]interface{}{ Args: map[string]interface{}{
"key": -12.25, "key": -12.25,
"foo": int64(-13), "foo": int64(-13),
@ -173,12 +173,12 @@ func TestParser_Parse(t *testing.T) {
// Parse with condition arguments. // Parse with condition arguments.
t.Run("WithCondition", func(t *testing.T) { t.Run("WithCondition", func(t *testing.T) {
q, err := pql.ParseString(`MyCall(key=foo, x == 12.25, y >= 100, z >< [4,8], m != null)`) q, err := pql.ParseString(`Row(key=foo, x == 12.25, y >= 100, z >< [4,8], m != null)`)
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} else if !reflect.DeepEqual(q.Calls[0], } else if !reflect.DeepEqual(q.Calls[0],
&pql.Call{ &pql.Call{
Name: "MyCall", Name: "Row",
Args: map[string]interface{}{ Args: map[string]interface{}{
"key": "foo", "key": "foo",
"x": &pql.Condition{Op: pql.EQ, Value: 12.25}, "x": &pql.Condition{Op: pql.EQ, Value: 12.25},

View file

@ -32,7 +32,7 @@ COND <- ( '><' { p.addBTWN() }
) )
conditional <- {p.startConditional()} condint condLT condfield condLT condint {p.endConditional()} conditional <- {p.startConditional()} condint condLT condfield condLT condint {p.endConditional()}
condint <- <'-'? [1-9] [0-9]* / '0'> sp {p.condAdd(buffer[begin:end])} condint <- < '-'? [0-9]* '.' [0-9]+ / '0' / '-'? [1-9] [0-9]* > sp {p.condAdd(buffer[begin:end])}
condLT <- <('<=' / '<')> sp {p.condAdd(buffer[begin:end])} condLT <- <('<=' / '<')> sp {p.condAdd(buffer[begin:end])}
condfield <- <fieldExpr> sp {p.condAdd(buffer[begin:end])} condfield <- <fieldExpr> sp {p.condAdd(buffer[begin:end])}
@ -55,7 +55,7 @@ item <- ( 'null' &(comma / sp close) { p.addVal(nil) }
doublequotedstring <- ( '\\"' / '\\\\' / [^"] )* doublequotedstring <- ( '\\"' / '\\\\' / [^"] )*
singlequotedstring <- ( '\\\'' / '\\\\' / [^'] )* singlequotedstring <- ( '\\\'' / '\\\\' / [^'] )*
fieldExpr <- [[A-Z]] ( [[A-Z]] / [0-9] / '_' / '-' )* fieldExpr <- ( [[A-Z]] / '_' ) ( [[A-Z]] / [0-9] / '_' / '-' )*
field <- <fieldExpr / reserved> { p.addField(buffer[begin:end]) } field <- <fieldExpr / reserved> { p.addField(buffer[begin:end]) }
reserved <- ('_row' / '_col' / '_start' / '_end' / '_timestamp' / '_field') reserved <- ('_row' / '_col' / '_start' / '_end' / '_timestamp' / '_field')
posfield <- <fieldExpr> { p.addPosStr("_field", buffer[begin:end]) } posfield <- <fieldExpr> { p.addPosStr("_field", buffer[begin:end]) }

File diff suppressed because it is too large Load diff

View file

@ -46,7 +46,7 @@ SetBit(Union(Zitmap(row==4), Intersect(Qitmap(blah>4), Ritmap(field="http://zoo9
t.Fatalf("Failed, got: %s", q) t.Fatalf("Failed, got: %s", q)
} }
_, err = ParseString("C(a=falsen0)") _, err = ParseString("Row(a=falsen0)")
if err != nil { if err != nil {
t.Fatalf("falsen0 should have been parsed as a string") t.Fatalf("falsen0 should have been parsed as a string")
} }
@ -109,15 +109,15 @@ func TestPEGWorking(t *testing.T) {
ncalls: 2}, ncalls: 2},
{ {
name: "SetWithArbCall", name: "SetWithArbCall",
input: "Set(1, a=4)Blerg(z=ha)", input: "Set(1, a=4)Row(z=ha)",
ncalls: 2}, ncalls: 2},
{ {
name: "SetArbSet", name: "SetArbSet",
input: "Set(1, a=4)Blerg(z=ha)Set(2, z=99)", input: "Set(1, a=4)Row(z=ha)Set(2, z=99)",
ncalls: 3}, ncalls: 3},
{ {
name: "ArbSetArb", name: "ArbSetArb",
input: "Arb(q=1, a=4)Set(1, z=9)Arb(z=99)", input: "Row(q=1, a=4)Set(1, z=9)Row(z=99)",
ncalls: 3}, ncalls: 3},
{ {
name: "SetStringArg", name: "SetStringArg",
@ -161,11 +161,11 @@ func TestPEGWorking(t *testing.T) {
ncalls: 1}, ncalls: 1},
{ {
name: "double quoted args", name: "double quoted args",
input: `B(a="zm''e")`, input: `Row(a="zm''e")`,
ncalls: 1}, ncalls: 1},
{ {
name: "single quoted args", name: "single quoted args",
input: `B(a='zm""e')`, input: `Row(a='zm""e')`,
ncalls: 1}, ncalls: 1},
{ {
name: "SetRowAttrs", name: "SetRowAttrs",
@ -320,7 +320,7 @@ func TestPEGErrors(t *testing.T) {
input: "Set(, 1, a=4)"}, input: "Set(, 1, a=4)"},
{ {
name: "StartinCommaArb", name: "StartinCommaArb",
input: "Zeeb(, a=4)"}, input: "Row(, a=4)"},
{ {
name: "SetRowAttrs0args", name: "SetRowAttrs0args",
input: "SetRowAttrs(blah, 9)"}, input: "SetRowAttrs(blah, 9)"},
@ -533,8 +533,8 @@ func TestPQLDeepEquality(t *testing.T) {
Name: "Row", Name: "Row",
Args: map[string]interface{}{ Args: map[string]interface{}{
"a": &Condition{ "a": &Condition{
Op: BETWEEN, Op: BTWN_LTE_LT,
Value: []interface{}{int64(4), int64(8)}, Value: []interface{}{int64(4), int64(9)},
}, },
}, },
}}, }},
@ -545,8 +545,8 @@ func TestPQLDeepEquality(t *testing.T) {
Name: "Row", Name: "Row",
Args: map[string]interface{}{ Args: map[string]interface{}{
"a": &Condition{ "a": &Condition{
Op: BETWEEN, Op: BTWN_LT_LT,
Value: []interface{}{int64(5), int64(8)}, Value: []interface{}{int64(4), int64(9)},
}, },
}, },
}}, }},
@ -569,8 +569,8 @@ func TestPQLDeepEquality(t *testing.T) {
Name: "Row", Name: "Row",
Args: map[string]interface{}{ Args: map[string]interface{}{
"a": &Condition{ "a": &Condition{
Op: BETWEEN, Op: BTWN_LT_LTE,
Value: []interface{}{int64(5), int64(9)}, Value: []interface{}{int64(4), int64(9)},
}, },
}, },
}}, }},
@ -585,11 +585,11 @@ func TestPQLDeepEquality(t *testing.T) {
}}, }},
{ {
name: "Weird dash", name: "Weird dash",
call: "Sum(field-=f)", call: "Count(dashy-=f)",
exp: &Call{ exp: &Call{
Name: "Sum", Name: "Count",
Args: map[string]interface{}{ Args: map[string]interface{}{
"field-": "f", "dashy-": "f",
}, },
}}, }},
{ {
@ -672,8 +672,8 @@ func TestPQLDeepEquality(t *testing.T) {
Name: "Row", Name: "Row",
Args: map[string]interface{}{ Args: map[string]interface{}{
"a": &Condition{ "a": &Condition{
Op: BETWEEN, Op: BTWN_LT_LT,
Value: []interface{}{int64(5), int64(8)}, Value: []interface{}{int64(4), int64(9)},
}, },
}, },
}, },

View file

@ -28,7 +28,17 @@ const (
LTE // <= LTE // <=
GT // > GT // >
GTE // >= GTE // >=
BETWEEN // ><
BETWEEN // >< (this is like a <= x <= b)
// not used in lexing/parsing, but so that the parser can signal
// to the executor how to treat the arguments. We used to just add
// 1 to the arguments if they were LT so the executor could assume
// it was always <=, <=, but then we needed to support
// floats/decimals and couldn't do that any more.
BTWN_LT_LTE // a < x <= b
BTWN_LTE_LT // a <= x < b
BTWN_LT_LT // a < x < b
) )
var tokens = [...]string{ var tokens = [...]string{

267
proto/interface.go Normal file
View file

@ -0,0 +1,267 @@
// Copyright 2017 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package pilosa
import (
"errors"
"fmt"
"strings"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
)
// StreamClient is an interface for a stream
// which can return a RowResponse sent to a
// stream via Send().
type StreamClient interface {
Recv() (*RowResponse, error)
}
// StreamServer is an interface for a stream
// which can accept a RowResponse to be later
// returned by the stream via Recv().
type StreamServer interface {
Send(*RowResponse) error
}
// EOF acts as an io.EOF encoded into a RowResponse.
var EOF *RowResponse = &RowResponse{
StatusError: &StatusError{
Code: 0,
Message: "EOF",
},
}
// Error is a helper function to create a RowResponse
// based on an error message. If the error is a grpc
// Status, then the status code is passed through.
func Error(err error) *RowResponse {
status, _ := status.FromError(err)
return &RowResponse{
StatusError: &StatusError{
Code: uint32(status.Code()),
Message: status.Err().Error(),
},
}
}
// ErrorWrap prepends a message to the existing status
// error message.
func ErrorWrap(err error, message string) *RowResponse {
status, _ := status.FromError(err)
return &RowResponse{
StatusError: &StatusError{
Code: uint32(status.Code()),
Message: message + ": " + status.Err().Error(),
},
}
}
// ErrorWrapf prepends a message to the existing status
// error message with the format specifier.
func ErrorWrapf(err error, format string, args ...interface{}) *RowResponse {
status, _ := status.FromError(err)
return &RowResponse{
StatusError: &StatusError{
Code: uint32(status.Code()),
Message: fmt.Sprintf(format, args...) + ": " + status.Err().Error(),
},
}
}
// ErrorCode is a helper function to create a RowResponse
// based on a grpc status code and an error message.
func ErrorCode(err error, c codes.Code) *RowResponse {
return &RowResponse{
StatusError: &StatusError{
Code: uint32(c),
Message: err.Error(),
},
}
}
// RowResponseSorter implements the sort interface for a
// provided []RowResponse based on the column index, type,
// and sort direction.
type RowResponseSorter struct {
colIdx []int
colDescending []bool
colType []string
rrs []*RowResponse
}
// NewRowResponseSorter return a new RowResponseSorter. It
// does input validation and returns an error if the inputs
// aren't compatible.
func NewRowResponseSorter(idxs []int, dirs []bool, typs []string, rrs []*RowResponse) (*RowResponseSorter, error) {
// Ensure the input slices are non-empty and equal size.
if len(idxs) == 0 {
return nil, errors.New("index list cannot be empty")
}
if len(dirs) != len(idxs) || len(typs) != len(idxs) {
return nil, errors.New("index, direction, and type lists must be the same size")
}
// Ensure the provided data types are supported by the sorter.
for i := range typs {
switch typs[i] {
case "[]uint64", "[]string", "bool", "float64", "int64", "string", "uint64":
// pass
default:
return nil, fmt.Errorf("unsupported data type: %s", typs[i])
}
}
// Ensure max(colIdx) is within size of rr.Columns.
if len(rrs) > 0 {
var maxColIdx int
for i := range idxs {
if idxs[i] > maxColIdx {
maxColIdx = idxs[i]
}
}
if maxColIdx >= len(rrs[0].Columns) {
return nil, fmt.Errorf("column index is out of range: %d", maxColIdx)
}
}
return &RowResponseSorter{
colIdx: idxs,
colDescending: dirs,
colType: typs,
rrs: rrs,
}, nil
}
func (r RowResponseSorter) Len() int { return len(r.rrs) }
func (r RowResponseSorter) Swap(i, j int) { r.rrs[i], r.rrs[j] = r.rrs[j], r.rrs[i] }
func (r RowResponseSorter) Less(i, j int) bool {
ri := r.rrs[i]
rj := r.rrs[j]
for i, idx := range r.colIdx {
coli := ri.Columns[idx]
colj := rj.Columns[idx]
var comp int
switch r.colType[i] {
case "[]uint64":
ai := coli.GetUint64ArrayVal().Vals
aj := colj.GetUint64ArrayVal().Vals
comp = func() int {
for ii := 0; ii < len(ai); ii++ {
if len(aj) == ii {
return 1
}
piv := ai[ii]
pjv := aj[ii]
if piv == pjv {
continue
} else if piv < pjv {
return -1
} else {
return 1
}
}
if len(aj) > len(ai) {
return -1
}
return 0
}()
case "[]string":
ai := coli.GetStringArrayVal().Vals
aj := colj.GetStringArrayVal().Vals
comp = func() int {
for ii := 0; ii < len(ai); ii++ {
if len(aj) == ii {
return 1
}
sComp := strings.Compare(ai[ii], aj[ii])
if sComp == 0 {
continue
} else {
return sComp
}
}
if len(aj) > len(ai) {
return -1
}
return 0
}()
case "bool":
bi := coli.GetBoolVal()
bj := colj.GetBoolVal()
if bi == bj {
comp = 0
} else if !bi && bj {
comp = -1
} else {
comp = 1
}
case "float64":
fi := coli.GetFloat64Val()
fj := colj.GetFloat64Val()
if fi == fj {
comp = 0
} else if fi < fj {
comp = -1
} else {
comp = 1
}
case "int64":
ni := coli.GetInt64Val()
nj := colj.GetInt64Val()
if ni == nj {
comp = 0
} else if ni < nj {
comp = -1
} else {
comp = 1
}
case "string":
comp = strings.Compare(coli.GetStringVal(), colj.GetStringVal())
case "uint64":
ni := coli.GetUint64Val()
nj := colj.GetUint64Val()
if ni == nj {
comp = 0
} else if ni < nj {
comp = -1
} else {
comp = 1
}
}
isDescending := r.colDescending[i]
switch comp {
case 0:
continue
case -1:
if isDescending {
return false
}
return true
case 1:
if isDescending {
return true
}
return false
}
}
return false
}

1050
proto/pilosa.pb.go Normal file

File diff suppressed because it is too large Load diff

64
proto/pilosa.proto Normal file
View file

@ -0,0 +1,64 @@
syntax = "proto3";
package pilosa;
message QueryPQLRequest {
string index = 1;
string pql = 2;
}
message StatusError{
uint32 Code = 1;
string Message = 2;
}
message RowResponse{
repeated ColumnInfo headers = 1;
repeated ColumnResponse columns = 2;
StatusError StatusError = 3;
}
message ColumnInfo {
string name = 1;
string datatype = 2;
}
message ColumnResponse{
oneof columnVal {
string stringVal = 1;
uint64 uint64Val = 2;
int64 int64Val = 3;
bool boolVal = 4;
bytes blobVal = 5;
Uint64Array uint64ArrayVal = 6;
StringArray stringArrayVal = 7;
double float64Val = 8;
}
}
message InspectRequest {
string index = 1;
IdsOrKeys columns = 2;
repeated string filterFields = 3;
uint64 limit = 4;
uint64 offset = 5;
}
message Uint64Array {
repeated uint64 vals = 1;
}
message StringArray {
repeated string vals = 1;
}
message IdsOrKeys {
oneof type {
Uint64Array ids = 1;
StringArray keys = 2;
}
}
service Pilosa {
rpc QueryPQL(QueryPQLRequest) returns (stream RowResponse) {};
rpc Inspect(InspectRequest) returns (stream RowResponse) {};
}

View file

@ -267,9 +267,6 @@ func (c *Container) Freeze() *Container {
if c.flags&flagFrozen != 0 { if c.flags&flagFrozen != 0 {
return c return c
} }
// unmapOrClone should unmap-in-place because the existing
// container isn't frozen (or we'd already have returned it).
c = c.unmapOrClone()
c.flags |= flagFrozen c.flags |= flagFrozen
return c return c
} }
@ -419,6 +416,48 @@ func (c *Container) bitmap() []uint64 {
return *(*[]uint64)(unsafe.Pointer(&reflect.SliceHeader{Data: uintptr(unsafe.Pointer(c.pointer)), Len: int(c.len), Cap: int(c.cap)})) return *(*[]uint64)(unsafe.Pointer(&reflect.SliceHeader{Data: uintptr(unsafe.Pointer(c.pointer)), Len: int(c.len), Cap: int(c.cap)}))
} }
// AsBitmap yields a 65k-bit bitmap, storing it in the target if a target
// is provided. The target should be zeroed, or this becomes an implicit
// union.
func (c *Container) AsBitmap(target []uint64) (out []uint64) {
if c.typeID == containerBitmap {
return c.bitmap()
}
// Reminder: len(nil) == 0.
if len(target) < 1024 {
out = make([]uint64, 1024)
} else {
out = target
for i := range out {
out[i] = 0
}
}
if c.typeID == containerArray {
a := c.array()
for _, v := range a {
out[v/64] |= 1 << (v % 64)
}
return out
}
if c.typeID == containerRun {
runs := c.runs()
for _, r := range runs {
splatRun(out, r)
}
return out
}
// in theory this shouldn't happen?
return out
}
func splatRun(into []uint64, from interval16) {
// TODO this can be ~64x faster for long runs by setting maxBitmap instead of single bits
//note v must be int or will overflow
for v := int(from.start); v <= int(from.last); v++ {
into[v/64] |= (uint64(1) << uint(v%64))
}
}
// setBitmap stores a set of uint64s as data. // setBitmap stores a set of uint64s as data.
func (c *Container) setBitmap(bitmap []uint64) { func (c *Container) setBitmap(bitmap []uint64) {
if c == nil || c.frozen() { if c == nil || c.frozen() {

View file

@ -212,7 +212,7 @@ func (sc *sliceContainers) Update(key uint64, fn func(*Container, bool) (*Contai
// don't expand the slice just to add a nil container, we // don't expand the slice just to add a nil container, we
// could return that anyway // could return that anyway
if write && nc != nil { if write && nc != nil {
sc.insertAt(key, nc, -i-1) sc.insertAt(key, nc, i)
} }
} }
} }

View file

@ -0,0 +1,19 @@
// Copyright 2019 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
// +build generationdebug
package roaring
const generationDebug = true

View file

@ -0,0 +1,19 @@
// Copyright 2019 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
// +build !generationdebug
package roaring
const generationDebug = false

View file

@ -21,6 +21,7 @@ import (
"hash/fnv" "hash/fnv"
"io" "io"
"math/bits" "math/bits"
"reflect"
"sort" "sort"
"unsafe" "unsafe"
@ -77,6 +78,46 @@ var containerTypeNames = map[byte]string{
var fullContainer = NewContainerRun([]interval16{{start: 0, last: maxContainerVal}}).Freeze() var fullContainer = NewContainerRun([]interval16{{start: 0, last: maxContainerVal}}).Freeze()
// AdvisoryError is used for the special case where we probably want to *report*
// an error reading a file, but don't want to actually count the file as not
// being read. For instance, a partial ops-log entry is *probably* harmless;
// we probably crashed while writing (?) and as such didn't report the write
// as successful. We hope.
type AdvisoryError interface {
error
AdvisoryOnly()
}
type advisoryError struct {
e error
}
func (a advisoryError) Error() string {
return a.e.Error()
}
// This marks the error as safe to ignore.
func (a advisoryError) AdvisoryOnly() {
}
type FileShouldBeTruncatedError interface {
AdvisoryError
SuggestedLength() int64
}
type fileShouldBeTruncatedError struct {
advisoryError
offset int64
}
func (f *fileShouldBeTruncatedError) SuggestedLength() int64 {
return f.offset
}
func newFileShouldBeTruncatedError(err error, offset int64) *fileShouldBeTruncatedError {
return &fileShouldBeTruncatedError{advisoryError: advisoryError{e: err}, offset: offset}
}
type Containers interface { type Containers interface {
// Get returns nil if the key does not exist. // Get returns nil if the key does not exist.
Get(key uint64) *Container Get(key uint64) *Container
@ -144,6 +185,7 @@ type ContainerIterator interface {
// Bitmap represents a roaring bitmap. // Bitmap represents a roaring bitmap.
type Bitmap struct { type Bitmap struct {
Containers Containers Containers Containers
Source Source
// User-defined flags. // User-defined flags.
Flags byte Flags byte
@ -218,6 +260,7 @@ func (b *Bitmap) Freeze() *Bitmap {
// Create a copy of the bitmap structure. // Create a copy of the bitmap structure.
other := &Bitmap{ other := &Bitmap{
Containers: b.Containers.Freeze(), Containers: b.Containers.Freeze(),
Source: b.Source,
} }
return other return other
@ -391,6 +434,13 @@ func (b *Bitmap) Min() (uint64, bool) {
return v, !eof return v, !eof
} }
// MinAt returns the lowest value in the bitmap at least equal to its argument.
// Second return value is true if containers exist in the bitmap.
func (b *Bitmap) MinAt(start uint64) (uint64, bool) {
v, eof := b.IteratorAt(start).Next()
return v, !eof
}
// Max returns the highest value in the bitmap. // Max returns the highest value in the bitmap.
// Returns zero if the bitmap is empty. // Returns zero if the bitmap is empty.
func (b *Bitmap) Max() uint64 { func (b *Bitmap) Max() uint64 {
@ -549,13 +599,21 @@ func (b *Bitmap) OffsetRange(offset, start, end uint64) *Bitmap {
hi0, hi1 := highbits(start), highbits(end) hi0, hi1 := highbits(start), highbits(end)
citer, _ := b.Containers.Iterator(hi0) citer, _ := b.Containers.Iterator(hi0)
other := NewSliceBitmap() other := NewSliceBitmap()
mappedAny := false
for citer.Next() { for citer.Next() {
k, c := citer.Value() k, c := citer.Value()
if k >= hi1 { if k >= hi1 {
break break
} }
if c.Mapped() {
mappedAny = true
}
other.Containers.Put(off+(k-hi0), c.Freeze()) other.Containers.Put(off+(k-hi0), c.Freeze())
} }
// if b.Source != nil && mappedAny {
if b.Source != nil && (generationDebug || mappedAny) {
other.Source = b.Source
}
return other return other
} }
@ -594,6 +652,7 @@ func (b *Bitmap) IntersectionCount(other *Bitmap) uint64 {
// Intersect returns the intersection of b and other. // Intersect returns the intersection of b and other.
func (b *Bitmap) Intersect(other *Bitmap) *Bitmap { func (b *Bitmap) Intersect(other *Bitmap) *Bitmap {
output := NewBitmap() output := NewBitmap()
usedB, usedOther := false, false
iiter, _ := b.Containers.Iterator(0) iiter, _ := b.Containers.Iterator(0)
jiter, _ := other.Containers.Iterator(0) jiter, _ := other.Containers.Iterator(0)
i, j := iiter.Next(), jiter.Next() i, j := iiter.Next(), jiter.Next()
@ -607,12 +666,27 @@ func (b *Bitmap) Intersect(other *Bitmap) *Bitmap {
j = jiter.Next() j = jiter.Next()
kj, cj = jiter.Value() kj, cj = jiter.Value()
} else { // ki == kj } else { // ki == kj
output.Containers.Put(ki, intersect(ci, cj)) newC := intersect(ci, cj)
if newC == ci {
usedB = true
}
if newC == cj {
usedOther = true
}
output.Containers.Put(ki, newC)
i, j = iiter.Next(), jiter.Next() i, j = iiter.Next(), jiter.Next()
ki, ci = iiter.Value() ki, ci = iiter.Value()
kj, cj = jiter.Value() kj, cj = jiter.Value()
} }
} }
switch {
case usedB && usedOther:
output.Source = MergeSources(b.Source, other.Source)
case usedB:
output.Source = b.Source
case usedOther:
output.Source = other.Source
}
return output return output
} }
@ -640,25 +714,43 @@ func (b *Bitmap) UnionInPlace(others ...*Bitmap) {
func (b *Bitmap) unionIntoTargetSingle(target *Bitmap, other *Bitmap) { func (b *Bitmap) unionIntoTargetSingle(target *Bitmap, other *Bitmap) {
iiter, _ := b.Containers.Iterator(0) iiter, _ := b.Containers.Iterator(0)
jiter, _ := other.Containers.Iterator(0) jiter, _ := other.Containers.Iterator(0)
usedB, usedOther := false, false
i, j := iiter.Next(), jiter.Next() i, j := iiter.Next(), jiter.Next()
ki, ci := iiter.Value() ki, ci := iiter.Value()
kj, cj := jiter.Value() kj, cj := jiter.Value()
for i || j { for i || j {
if i && (!j || ki < kj) { if i && (!j || ki < kj) {
target.Containers.Put(ki, ci.Freeze()) target.Containers.Put(ki, ci.Freeze())
usedB = true
i = iiter.Next() i = iiter.Next()
ki, ci = iiter.Value() ki, ci = iiter.Value()
} else if j && (!i || ki > kj) { } else if j && (!i || ki > kj) {
target.Containers.Put(kj, cj.Freeze()) target.Containers.Put(kj, cj.Freeze())
usedOther = true
j = jiter.Next() j = jiter.Next()
kj, cj = jiter.Value() kj, cj = jiter.Value()
} else { // ki == kj } else { // ki == kj
target.Containers.Put(ki, union(ci, cj)) newC := union(ci, cj)
target.Containers.Put(ki, newC)
if newC == ci {
usedB = true
}
if newC == cj {
usedOther = true
}
i, j = iiter.Next(), jiter.Next() i, j = iiter.Next(), jiter.Next()
ki, ci = iiter.Value() ki, ci = iiter.Value()
kj, cj = jiter.Value() kj, cj = jiter.Value()
} }
} }
switch {
case usedB && usedOther:
target.Source = MergeSources(b.Source, other.Source)
case usedB:
target.Source = b.Source
case usedOther:
target.Source = other.Source
}
} }
// unionInPlace stores the union of b and others into b. The others will // unionInPlace stores the union of b and others into b. The others will
@ -752,7 +844,14 @@ func (b *Bitmap) unionInPlace(others ...*Bitmap) {
bitmapIters = make(handledIters, 0, requiredSliceSize) bitmapIters = make(handledIters, 0, requiredSliceSize)
} }
var sources []Source
if b.Source != nil {
sources = append(sources, b.Source)
}
for _, other := range others { for _, other := range others {
if other.Source != nil {
sources = append(sources, other.Source)
}
otherIter, _ := other.Containers.Iterator(0) otherIter, _ := other.Containers.Iterator(0)
if otherIter.Next() { if otherIter.Next() {
bitmapIters = append(bitmapIters, handledIter{ bitmapIters = append(bitmapIters, handledIter{
@ -762,6 +861,8 @@ func (b *Bitmap) unionInPlace(others ...*Bitmap) {
}) })
} }
} }
// new bitmap might have containers from any of those bitmaps in it
b.Source = MergeSources(sources...)
// Loop until we've exhausted every iter. // Loop until we've exhausted every iter.
hasNext := true hasNext := true
@ -890,6 +991,7 @@ func (b *Bitmap) unionInPlace(others ...*Bitmap) {
// Difference returns the difference of b and other. // Difference returns the difference of b and other.
func (b *Bitmap) Difference(other *Bitmap) *Bitmap { func (b *Bitmap) Difference(other *Bitmap) *Bitmap {
output := NewBitmap() output := NewBitmap()
output.Source = b.Source
iiter, _ := b.Containers.Iterator(0) iiter, _ := b.Containers.Iterator(0)
jiter, _ := other.Containers.Iterator(0) jiter, _ := other.Containers.Iterator(0)
@ -917,6 +1019,9 @@ func (b *Bitmap) Difference(other *Bitmap) *Bitmap {
// Xor returns the bitwise exclusive or of b and other. // Xor returns the bitwise exclusive or of b and other.
func (b *Bitmap) Xor(other *Bitmap) *Bitmap { func (b *Bitmap) Xor(other *Bitmap) *Bitmap {
output := NewBitmap() output := NewBitmap()
// Xor can end up with containers from either parent if the other
// had no container or an empty container.
output.Source = MergeSources(b.Source, other.Source)
iiter, _ := b.Containers.Iterator(0) iiter, _ := b.Containers.Iterator(0)
jiter, _ := other.Containers.Iterator(0) jiter, _ := other.Containers.Iterator(0)
@ -1433,7 +1538,11 @@ func (b *Bitmap) RemapRoaringStorage(data []byte) (mappedAny bool, returnErr err
var itrPointer *uint16 var itrPointer *uint16
var itrErr error var itrErr error
if data != nil { // If we got no data, we don't want to do the actual mapping, just
// the unmapping. If preferMapping is false, we also don't want to
// map to the data. We still need to do the UpdateEvery loop, we
// just won't have an iterator for it.
if data != nil && b.preferMapping {
itr, err = newRoaringIterator(data) itr, err = newRoaringIterator(data)
} }
// don't return early: we still have to do the unmapping // don't return early: we still have to do the unmapping
@ -1617,6 +1726,12 @@ func (b *Bitmap) Iterator() *Iterator {
return itr return itr
} }
func (b *Bitmap) IteratorAt(start uint64) *Iterator {
itr := &Iterator{bitmap: b}
itr.Seek(start)
return itr
}
// Ops returns the number of write ops the bitmap is aware of in its ops // Ops returns the number of write ops the bitmap is aware of in its ops
// log, and their total bit count. // log, and their total bit count.
func (b *Bitmap) Ops() (ops int, opN int) { func (b *Bitmap) Ops() (ops int, opN int) {
@ -1629,6 +1744,167 @@ func (b *Bitmap) SetOps(ops int, opN int) {
b.ops, b.opN = ops, opN b.ops, b.opN = ops, opN
} }
// RoaringToBitmaps yields a series of bitmaps with specified shard
// keys, based on a single roaring file, with splits at multiples of
// shardWidth, which should be a multiple of container size.
func RoaringToBitmaps(data []byte, shardWidth uint64) ([]*Bitmap, []uint64) {
if data == nil {
return nil, nil
}
var itr roaringIterator
var itrKey uint64
var itrCType byte
var itrN int
var itrLen int
var itrPointer *uint16
var itrErr error
currentShard := ^uint64(0)
var currentBitmap *Bitmap
var bitmaps []*Bitmap
var shards []uint64
keysPerShard := shardWidth >> 16
itr, err := newRoaringIterator(data)
if err != nil || itr == nil {
return nil, nil
}
itrKey, itrCType, itrN, itrLen, itrPointer, itrErr = itr.Next()
for itrErr == nil {
newC := &Container{
typeID: itrCType,
n: int32(itrN),
len: int32(itrLen),
cap: int32(itrLen),
pointer: itrPointer,
flags: flagMapped,
}
shard := itrKey / keysPerShard
if shard != currentShard {
if currentBitmap != nil {
bitmaps = append(bitmaps, currentBitmap)
shards = append(shards, currentShard)
}
currentBitmap = NewFileBitmap()
currentShard = shard
}
currentBitmap.Containers.Put(itrKey, newC)
itrKey, itrCType, itrN, itrLen, itrPointer, itrErr = itr.Next()
}
if currentBitmap != nil {
bitmaps = append(bitmaps, currentBitmap)
shards = append(shards, currentShard)
}
// we don't support ops logs for this
return bitmaps, shards
}
// BitmapsToRoaring renders a series of non-overlapping bitmaps as a
// unified roaring file.
func BitmapsToRoaring(bitmaps []*Bitmap) []byte {
count := int64(0)
size := int64(0)
for i, bm := range bitmaps {
c, s := bm.roaringSize()
// skip this bitmap during the next pass, since it's empty
if c == 0 {
bitmaps[i] = nil
continue
}
count += c
size += s
}
if count == 0 {
return nil
}
// we have count headers, which need 12 bytes, plus a magic number,
// plus offsets (4 bytes per container), plus size bytes of data to
// write.
out := make([]byte, headerBaseSize+(12*count)+(4*count)+size)
binary.LittleEndian.PutUint16(out[0:2], uint16(MagicNumber))
out[3] = byte(storageVersion)
binary.LittleEndian.PutUint32(out[4:8], uint32(count))
headerEnd := 8 + (12 * count)
offsetEnd := headerEnd + (4 * count)
headers := out[8:headerEnd]
offsets := out[headerEnd:offsetEnd]
data := out[offsetEnd:]
headerOffset := 0
offsetOffset := 0
dataOffset := 0
prevKey := uint64(0)
for _, bm := range bitmaps {
if bm == nil {
continue
}
citer, _ := bm.Containers.Iterator(0)
for citer.Next() {
k, c := citer.Value()
n := c.N()
if n == 0 {
continue
}
if roaringParanoia {
if k < prevKey {
panic("unsorted keys in multiple-bitmap roaring conversion")
}
}
// place header at header offset, and data at data
// offset
header := headers[headerOffset : headerOffset+12]
offset := offsets[offsetOffset : offsetOffset+4]
headerOffset += 12
offsetOffset += 4
binary.LittleEndian.PutUint64(header[0:8], k)
binary.LittleEndian.PutUint16(header[8:10], uint16(c.typeID))
binary.LittleEndian.PutUint16(header[10:12], uint16(n-1))
binary.LittleEndian.PutUint32(offset[0:4], uint32(dataOffset+int(offsetEnd)))
nextData := data[dataOffset:]
switch c.typeID {
case containerArray:
asUint16 := *(*[]uint16)(unsafe.Pointer(&reflect.SliceHeader{Data: uintptr(unsafe.Pointer(&nextData[0])), Len: int(c.len), Cap: int(c.len)}))
copy(asUint16, c.array())
dataOffset += 2 * int(c.len)
case containerBitmap:
asUint64 := *(*[]uint64)(unsafe.Pointer(&reflect.SliceHeader{Data: uintptr(unsafe.Pointer(&nextData[0])), Len: 1024, Cap: 1024}))
copy(asUint64, c.bitmap())
dataOffset += 8192
case containerRun:
asInterval16 := *(*[]interval16)(unsafe.Pointer(&reflect.SliceHeader{Data: uintptr(unsafe.Pointer(&nextData[2])), Len: int(c.len), Cap: int(c.len)}))
copy(asInterval16, c.runs())
binary.LittleEndian.PutUint16(nextData[0:2], uint16(c.len))
dataOffset += int(4*c.len) + 2
}
}
}
return out
}
// roaringSize yields the count of non-empty containers, and the size
// of the storage *only* -- not the headers.
func (b *Bitmap) roaringSize() (int64, int64) {
count := int64(0)
size := int64(0)
citer, _ := b.Containers.Iterator(0)
for citer.Next() {
_, c := citer.Value()
if c.N() == 0 {
continue
}
count++
switch c.typeID {
case containerArray:
size += 2 * int64(c.N())
case containerBitmap:
size += 8192
case containerRun:
// 2 bytes for the count of runs, plus 4 bytes per run
size += 2 + (4 * int64(c.len))
}
}
return count, size
}
// Info returns stats for the bitmap. // Info returns stats for the bitmap.
func (b *Bitmap) Info() bitmapInfo { func (b *Bitmap) Info() bitmapInfo {
info := bitmapInfo{ info := bitmapInfo{
@ -3271,9 +3547,9 @@ func intersectRunRun(a, b *Container) *Container {
output.setN(n) output.setN(n)
runs := output.runs() runs := output.runs()
if n < ArrayMaxSize && int32(len(runs)) > n/2 { if n < ArrayMaxSize && int32(len(runs)) > n/2 {
output.runToArray() output = output.runToArray()
} else if len(runs) > runMaxSize { } else if len(runs) > runMaxSize {
output.runToBitmap() output = output.runToBitmap()
} }
return output return output
} }
@ -3412,41 +3688,43 @@ func union(a, b *Container) *Container {
func unionArrayArray(a, b *Container) *Container { func unionArrayArray(a, b *Container) *Container {
statsHit("union/ArrayArray") statsHit("union/ArrayArray")
aa, ab := a.array(), b.array() if a.N() == 0 {
na, nb := len(aa), len(ab) return b
output := make([]uint16, na+nb)
n := 0
for i, j := 0, 0; ; {
if i >= na && j >= nb {
break
} else if i < na && j >= nb {
output[n] = aa[i]
n++
i++
continue
} else if i >= na && j < nb {
output[n] = ab[j]
n++
j++
continue
} }
if b.N() == 0 {
va, vb := aa[i], ab[j] return a
}
s1, s2 := a.array(), b.array()
n1, n2 := len(s1), len(s2)
output := make([]uint16, 0, n1+n2)
i, j := 0, 0
for {
va, vb := s1[i], s2[j]
if va < vb { if va < vb {
output[n] = va output = append(output, va)
n++
i++ i++
} else if va > vb { } else if va > vb {
output[n] = vb output = append(output, vb)
n++
j++ j++
} else { } else {
output[n] = va output = append(output, va)
n++ i++
i, j = i+1, j+1 j++
}
// It's possible we hit the ends at the same time,
// in which case the append will copy 0 items. This
// is cheaper than performing a separate conditional
// check every time...
if j >= n2 {
output = append(output, s1[i:]...)
break
}
if i >= n1 {
output = append(output, s2[j:]...)
break
} }
} }
return NewContainerArray(output[:n]) return NewContainerArray(output)
} }
// unionArrayArrayInPlace does what it sounds like -- tries to combine // unionArrayArrayInPlace does what it sounds like -- tries to combine
@ -3454,47 +3732,56 @@ func unionArrayArray(a, b *Container) *Container {
// of a good array size, so it could be up to twice that size, temporarily. // of a good array size, so it could be up to twice that size, temporarily.
func unionArrayArrayInPlace(a, b *Container) *Container { func unionArrayArrayInPlace(a, b *Container) *Container {
statsHit("union/ArrayArrayInPlace") statsHit("union/ArrayArrayInPlace")
aa, ab := a.array(), b.array() if a.N() == 0 {
na, nb := len(aa), len(ab) if b.N() != 0 {
output := make([]uint16, na+nb) // for InPlace, we actually want to ensure that
outN := 0 // we update a, as long as it's not frozen.
for i, j := 0, 0; ; { a = a.Thaw()
if i >= na && j >= nb { a.setArray(b.array())
break return a.optimize()
} else if i < na && j >= nb {
copy(output[outN:], aa[i:])
outN += na - i
break
} else if i >= na && j < nb {
copy(output[outN:], ab[j:])
outN += nb - j
break
} }
return a
va, vb := aa[i], ab[j] }
if b.N() == 0 {
return a
}
s1, s2 := a.array(), b.array()
n1, n2 := len(s1), len(s2)
output := make([]uint16, 0, n1+n2)
i, j := 0, 0
for {
va, vb := s1[i], s2[j]
if va < vb { if va < vb {
output[outN] = va output = append(output, va)
outN++
i++ i++
} else if va > vb { } else if va > vb {
output[outN] = vb output = append(output, vb)
outN++
j++ j++
} else { } else {
output[outN] = va output = append(output, va)
outN++
i++ i++
j++ j++
} }
// It's possible we hit the ends at the same time,
// in which case the append will copy 0 items. This
// is cheaper than performing a separate conditional
// check every time...
if j >= n2 {
output = append(output, s1[i:]...)
break
}
if i >= n1 {
output = append(output, s2[j:]...)
break
}
} }
// a union can't omit anything that was previously in a, so if // a union can't omit anything that was previously in a, so if
// the output is the same length, nothing changed. // the output is the same length, nothing changed.
if len(output) != int(a.N()) { if len(output) != int(a.N()) {
a = a.Thaw() a = a.Thaw()
a.setArray(output[:outN]) a.setArray(output)
a = a.optimize()
} }
return a return a.optimize()
} }
// unionArrayRun optimistically assumes that the result will be a run container, // unionArrayRun optimistically assumes that the result will be a run container,

98
roaring/source.go Normal file
View file

@ -0,0 +1,98 @@
// Copyright 2017 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package roaring
import (
"strings"
)
// A Source represents the source a given bitmap gets its data from,
// such as a memory-mapped file. When combining bitmaps, we might
// track them together in a single combined-source of some sort.
type Source interface {
ID() string
Dead() bool
}
// MergeSources combines sources. If you have two bitmaps, and you're
// combining them, then the combination's source is a combination of
// those two sources.
func MergeSources(sources ...Source) Source {
sourceCount := 0
totalCount := 0
var lastSource Source
for _, s := range sources {
if s == nil {
continue
}
lastSource = s
if s, ok := s.(combinedSource); ok {
sourceCount++
totalCount += len(s)
} else {
sourceCount++
totalCount++
}
}
// if there's no sources (this includes all sources being
// empty combinedSources), we don't have a source.
if totalCount == 0 {
return nil
}
// if there's exactly one source, combined or otherwise, that's
// fine, we'll just return it.
if sourceCount == 1 {
return lastSource
}
// make a new combinedSource, flattening any combinedSources
// already present.
newSources := make([]Source, 0, totalCount)
for _, s := range sources {
if s == nil {
continue
}
if s, ok := s.(combinedSource); ok {
newSources = append(newSources, s...)
} else {
newSources = append(newSources, s)
}
}
return combinedSource(newSources)
}
// SetSource tells the bitmap what source to associate with new things it
// creates. This is possibly logically incorrect.
func (b *Bitmap) SetSource(s Source) {
b.Source = s
}
type combinedSource []Source
func (c combinedSource) ID() string {
ids := make([]string, len(c))
for i := range c {
ids[i] = c[i].ID()
}
return strings.Join(ids, ",")
}
func (c combinedSource) Dead() bool {
for i := range c {
if c[i].Dead() {
return true
}
}
return false
}

View file

@ -30,7 +30,9 @@ func (b *Bitmap) UnmarshalBinary(data []byte) error {
return nil return nil
} }
statsHit("Bitmap/UnmarshalBinary") statsHit("Bitmap/UnmarshalBinary")
b.opN = 0 // reset opN since we're reading new data. // reset ops/opN since we're reading new data.
b.ops = 0
b.opN = 0
fileMagic := uint32(binary.LittleEndian.Uint16(data[0:2])) fileMagic := uint32(binary.LittleEndian.Uint16(data[0:2]))
if fileMagic == MagicNumber { // if pilosa roaring if fileMagic == MagicNumber { // if pilosa roaring
return errors.Wrap(b.unmarshalPilosaRoaring(data), "unmarshaling as pilosa roaring") return errors.Wrap(b.unmarshalPilosaRoaring(data), "unmarshaling as pilosa roaring")
@ -205,15 +207,15 @@ func (b *Bitmap) unmarshalPilosaRoaring(data []byte) error {
// Unmarshal the op and apply it. // Unmarshal the op and apply it.
var opr op var opr op
if err := opr.UnmarshalBinary(buf); err != nil { if err := opr.UnmarshalBinary(buf); err != nil {
// FIXME(benbjohnson): return error with position so file can be trimmed. return newFileShouldBeTruncatedError(err, int64(opsOffset))
return err
} }
opr.apply(b) opr.apply(b)
// Increase the op count. // Increase the op count.
b.ops++ b.ops++
b.opN += opr.count() b.opN += opr.count()
opsOffset += opr.size()
// Move the buffer forward. // Move the buffer forward.
buf = buf[opr.size():] buf = data[opsOffset:]
} }
return nil return nil

204
row.go
View file

@ -18,6 +18,7 @@ import (
"encoding/json" "encoding/json"
"sort" "sort"
"github.com/molecula/ext"
"github.com/pilosa/pilosa/v2/roaring" "github.com/pilosa/pilosa/v2/roaring"
"github.com/pkg/errors" "github.com/pkg/errors"
) )
@ -43,6 +44,55 @@ func NewRow(columns ...uint64) *Row {
return r return r
} }
// NewRowFromBitmap divides a bitmap into rows, which it now calls shards. This
// transposes; data that was in any shard for Row 0 is now considered shard 0,
// etcetera.
func NewRowFromBitmap(b *roaring.Bitmap) *Row {
r := &Row{}
if b == nil {
return r
}
rowNum := uint64(0)
for col, ok := b.MinAt(rowNum * ShardWidth); ok; col, ok = b.MinAt(rowNum * ShardWidth) {
rowNum = col / ShardWidth
seg := rowSegment{
shard: rowNum,
data: b.OffsetRange(rowNum*ShardWidth, rowNum*ShardWidth, (rowNum+1)*ShardWidth),
writable: true,
}
seg.n = seg.data.Count()
r.segments = append(r.segments, seg)
rowNum++
}
return r
}
// NewRowFromRoaring parses a roaring data file as a row, dividing it into
// bitmaps and rowSegments based on shard width.
func NewRowFromRoaring(data []byte) *Row {
bitmaps, shards := roaring.RoaringToBitmaps(data, ShardWidth)
r := &Row{segments: make([]rowSegment, len(bitmaps))}
for i := range bitmaps {
segment := rowSegment{
shard: shards[i],
data: bitmaps[i],
writable: false,
n: bitmaps[i].Count(),
}
r.segments[i] = segment
}
return r
}
// Roaring returns the row treated as a unified roaring bitmap.
func (r *Row) Roaring() []byte {
bitmaps := make([]*roaring.Bitmap, len(r.segments))
for i := range r.segments {
bitmaps[i] = r.segments[i].data
}
return roaring.BitmapsToRoaring(bitmaps)
}
// IsEmpty returns true if the row doesn't contain any set bits. // IsEmpty returns true if the row doesn't contain any set bits.
func (r *Row) IsEmpty() bool { func (r *Row) IsEmpty() bool {
if len(r.segments) == 0 { if len(r.segments) == 0 {
@ -194,6 +244,65 @@ func (r *Row) Union(others ...*Row) *Row {
return &Row{segments: output} return &Row{segments: output}
} }
// GenericBinaryOp returns the output of a generic op on r and other.
func (r *Row) GenericBinaryOp(op ext.GenericBitmapOpBitmap, other *Row, args map[string]interface{}) *Row {
var segments []rowSegment
itr := newMergeSegmentIterator(r.segments, other.segments)
for s0, s1 := itr.next(); s0 != nil || s1 != nil; s0, s1 = itr.next() {
if s1 == nil {
segments = append(segments, *s0)
continue
} else if s0 == nil {
segments = append(segments, *s1)
continue
}
segments = append(segments, *s0.GenericBinaryOp(op, s1, args))
}
return &Row{segments: segments}
}
// GenericNaryOp returns the output of an nary op on r and others.
func (r *Row) GenericNaryOp(op ext.GenericBitmapOpBitmap, others []*Row, args map[string]interface{}) *Row {
segments := make([][]rowSegment, 0, len(others)+1)
if len(r.segments) > 0 {
segments = append(segments, r.segments)
}
nextSegs := make([][]rowSegment, 0, len(others)+1)
toProcess := make([]*rowSegment, 0, len(others)+1)
var output []rowSegment
for _, other := range others {
if len(other.segments) > 0 {
segments = append(segments, other.segments)
}
}
for len(segments) > 0 {
shard := segments[0][0].shard
for _, segs := range segments {
if segs[0].shard < shard {
shard = segs[0].shard
}
}
nextSegs = nextSegs[:0]
toProcess := toProcess[:0]
for _, segs := range segments {
if segs[0].shard == shard {
toProcess = append(toProcess, &segs[0])
segs = segs[1:]
}
if len(segs) > 0 {
nextSegs = append(nextSegs, segs)
}
}
// at this point, "toProcess" is a list of all the segments
// sharing the lowest ID, and nextSegs is a list of all the others.
// Swap the segment lists (so we don't have to reallocate it)
segments, nextSegs = nextSegs, segments
output = append(output, *toProcess[0].GenericNaryOp(op, toProcess[1:], args))
}
return &Row{segments: output}
}
// Difference returns the diff of r and other. // Difference returns the diff of r and other.
func (r *Row) Difference(other *Row) *Row { func (r *Row) Difference(other *Row) *Row {
var segments []rowSegment var segments []rowSegment
@ -212,6 +321,17 @@ func (r *Row) Difference(other *Row) *Row {
return &Row{segments: segments} return &Row{segments: segments}
} }
// GenericUnary returns the results of a generic op on r.
func (r *Row) GenericUnaryOp(op ext.GenericBitmapOpBitmap, args map[string]interface{}) *Row {
work := r
var segments []rowSegment
for _, segment := range work.segments {
opped := segment.GenericUnaryOp(op, args)
segments = append(segments, *opped)
}
return &Row{segments: segments}
}
// Shift returns the bitwise shift of r by n bits. // Shift returns the bitwise shift of r by n bits.
// Currently only positive shift values are supported. // Currently only positive shift values are supported.
func (r *Row) Shift(n int64) (*Row, error) { func (r *Row) Shift(n int64) (*Row, error) {
@ -299,6 +419,15 @@ func (r *Row) Count() uint64 {
return n return n
} }
// GenericCount applies an op to lots of things.
func (r *Row) GenericCount(op ext.BitmapOpUnaryCount, args map[string]interface{}) uint64 {
var n int64
for i := range r.segments {
n += op([]ext.Bitmap{WrapBitmap(r.segments[i].data)}, args)
}
return uint64(n)
}
// MarshalJSON returns a JSON-encoded byte slice of r. // MarshalJSON returns a JSON-encoded byte slice of r.
func (r *Row) MarshalJSON() ([]byte, error) { func (r *Row) MarshalJSON() ([]byte, error) {
var o struct { var o struct {
@ -326,6 +455,21 @@ func (r *Row) Columns() []uint64 {
return a return a
} }
// Includes returns true if the row contains the given column.
func (r *Row) Includes(col uint64) bool {
// TODO: improve the efficiency of this method by
// performing the column filter at the bitmap level
// rather than iterating through the results here.
for i := range r.segments {
for _, c := range r.segments[i].Columns() {
if c == col {
return true
}
}
}
return false
}
// rowSegment holds a subset of a row. // rowSegment holds a subset of a row.
// This could point to a mmapped roaring bitmap or an in-memory bitmap. The // This could point to a mmapped roaring bitmap or an in-memory bitmap. The
// width of the segment will always match the shard width. // width of the segment will always match the shard width.
@ -344,9 +488,20 @@ type rowSegment struct {
} }
func (s *rowSegment) Freeze() { func (s *rowSegment) Freeze() {
s.data.Freeze() s.data = s.data.Freeze()
} }
/*
// Raw returns the row segment as a byte slice.
// It may be used by the gRPC server to deliver results
// as a roaring bitmap instead of a stream of RowResults.
func (s *rowSegment) Raw() (uint64, []byte) {
var buf bytes.Buffer
s.data.WriteTo(&buf)
return s.shard, buf.Bytes()
}
*/
// Merge adds chunks from other to s. // Merge adds chunks from other to s.
// Chunks in s are overwritten if they exist in other. // Chunks in s are overwritten if they exist in other.
func (s *rowSegment) Merge(other *rowSegment) { func (s *rowSegment) Merge(other *rowSegment) {
@ -366,7 +521,7 @@ func (s *rowSegment) IntersectionCount(other *rowSegment) uint64 {
// Intersect returns the itersection of s and other. // Intersect returns the itersection of s and other.
func (s *rowSegment) Intersect(other *rowSegment) *rowSegment { func (s *rowSegment) Intersect(other *rowSegment) *rowSegment {
data := s.data.Intersect(other.data) data := s.data.Intersect(other.data)
data.Freeze() data = data.Freeze()
return &rowSegment{ return &rowSegment{
data: data, data: data,
@ -393,10 +548,37 @@ func (s *rowSegment) Union(others ...*rowSegment) *rowSegment {
} }
} }
// GenericOp performs a generic op on s and other
func (s *rowSegment) GenericBinaryOp(op ext.GenericBitmapOpBitmap, other *rowSegment, args map[string]interface{}) *rowSegment {
data := op([]ext.Bitmap{WrapBitmap(s.data), WrapBitmap(other.data)}, args)
return &rowSegment{
data: UnwrapBitmap(data),
shard: s.shard,
n: data.Count(),
}
}
// GenericOp performs a generic op on s and others
func (s *rowSegment) GenericNaryOp(op ext.GenericBitmapOpBitmap, others []*rowSegment, args map[string]interface{}) *rowSegment {
bitmaps := make([]ext.Bitmap, len(others)+1)
bitmaps[0] = WrapBitmap(s.data)
for i, seg := range others {
bitmaps[i+1] = WrapBitmap(seg.data)
}
data := op(bitmaps, args)
return &rowSegment{
data: UnwrapBitmap(data),
shard: s.shard,
n: data.Count(),
}
}
// Difference returns the diff of s and other. // Difference returns the diff of s and other.
func (s *rowSegment) Difference(other *rowSegment) *rowSegment { func (s *rowSegment) Difference(other *rowSegment) *rowSegment {
data := s.data.Difference(other.data) data := s.data.Difference(other.data)
data.Freeze() data = data.Freeze()
return &rowSegment{ return &rowSegment{
data: data, data: data,
@ -409,7 +591,7 @@ func (s *rowSegment) Difference(other *rowSegment) *rowSegment {
// Xor returns the xor of s and other. // Xor returns the xor of s and other.
func (s *rowSegment) Xor(other *rowSegment) *rowSegment { func (s *rowSegment) Xor(other *rowSegment) *rowSegment {
data := s.data.Xor(other.data) data := s.data.Xor(other.data)
data.Freeze() data = data.Freeze()
return &rowSegment{ return &rowSegment{
data: data, data: data,
@ -426,7 +608,7 @@ func (s *rowSegment) Shift() (*rowSegment, error) {
if err != nil { if err != nil {
return nil, errors.Wrap(err, "shifting roaring data") return nil, errors.Wrap(err, "shifting roaring data")
} }
data.Freeze() data = data.Freeze()
return &rowSegment{ return &rowSegment{
data: data, data: data,
@ -436,6 +618,18 @@ func (s *rowSegment) Shift() (*rowSegment, error) {
}, nil }, nil
} }
// GenericUnary returns s subject to op.
func (s *rowSegment) GenericUnaryOp(op ext.GenericBitmapOpBitmap, args map[string]interface{}) *rowSegment {
//TODO deal with overflow
data := UnwrapBitmap(op([]ext.Bitmap{WrapBitmap(s.data)}, args))
return &rowSegment{
data: data,
shard: s.shard,
n: data.Count(),
}
}
// SetBit sets the i-th column of the row. // SetBit sets the i-th column of the row.
func (s *rowSegment) SetBit(i uint64) (changed bool) { func (s *rowSegment) SetBit(i uint64) (changed bool) {
s.ensureWritable() s.ensureWritable()

View file

@ -123,5 +123,18 @@ func TestRow_IsEmpty(t *testing.T) {
if !res.IsEmpty() { if !res.IsEmpty() {
t.Fatal("Result Should Be Empty\n") t.Fatal("Result Should Be Empty\n")
} }
}
func TestRow_Includes(t *testing.T) {
row := pilosa.NewRow(0, 2*ShardWidth)
if !row.Includes(0) {
t.Fatal("row should include 0")
}
if row.Includes(1) {
t.Fatal("row should not include 1")
}
if !row.Includes(2 * ShardWidth) {
t.Fatalf("row should include %d", 2*ShardWidth)
}
} }

View file

@ -27,7 +27,11 @@ import (
"sync" "sync"
"time" "time"
"github.com/molecula/ext"
// extensions pulls in some extensions depending on build tags
_ "github.com/pilosa/pilosa/v2/extensions"
"github.com/pilosa/pilosa/v2/logger" "github.com/pilosa/pilosa/v2/logger"
"github.com/pilosa/pilosa/v2/pql"
"github.com/pilosa/pilosa/v2/roaring" "github.com/pilosa/pilosa/v2/roaring"
"github.com/pilosa/pilosa/v2/stats" "github.com/pilosa/pilosa/v2/stats"
"github.com/pkg/errors" "github.com/pkg/errors"
@ -57,6 +61,7 @@ type Server struct { // nolint: maligned
hosts []string hosts []string
clusterDisabled bool clusterDisabled bool
serializer Serializer serializer Serializer
extensions []*ext.ExtensionInfo
// External // External
systemInfo SystemInfo systemInfo SystemInfo
@ -335,7 +340,6 @@ func NewServer(opts ...ServerOption) (*Server, error) {
if err != nil { if err != nil {
return nil, err return nil, err
} }
s.holder.Path = path s.holder.Path = path
// s.holder.translateFile.Path = filepath.Join(path, ".keys") // s.holder.translateFile.Path = filepath.Join(path, ".keys")
s.holder.Logger = s.logger s.holder.Logger = s.logger
@ -376,6 +380,30 @@ func NewServer(opts ...ServerOption) (*Server, error) {
s.cluster.broadcaster = s s.cluster.broadcaster = s
s.cluster.maxWritesPerRequest = s.maxWritesPerRequest s.cluster.maxWritesPerRequest = s.maxWritesPerRequest
s.holder.broadcaster = s s.holder.broadcaster = s
err = s.loadAllExtensions()
if err != nil {
s.logger.Printf("not all plugins loaded successfully")
}
if len(s.extensions) > 0 {
s.logger.Printf("loaded extensions:")
for _, ext := range s.extensions {
if ext == nil {
s.logger.Printf(" inexplicably, a nil extension?!?")
continue
}
s.logger.Printf(" %s %s: %s", ext.Name, ext.Version, ext.Description)
if ext.License != "" {
s.logger.Printf(" License: %s", ext.License)
}
if len(ext.BitmapOps) > 0 {
opList := make([]string, len(ext.BitmapOps))
for i := range ext.BitmapOps {
opList[i] = ext.BitmapOps[i].Name
}
s.logger.Printf(" Ops: %s", strings.Join(opList, ", "))
}
}
}
err = s.cluster.setup() err = s.cluster.setup()
if err != nil { if err != nil {
@ -389,6 +417,56 @@ func (s *Server) InternalClient() InternalClient {
return s.defaultClient return s.defaultClient
} }
// loadNewExtensions loads extensions that have been
// registered since the last call to loadNewExtensions.
func (s *Server) loadNewExtensions() error { //nolint:unused
return s.loadExtensions(ext.NewExtensions())
}
// loadAllExtensions loads all extensions.
func (s *Server) loadAllExtensions() error {
return s.loadExtensions(ext.AllExtensions())
}
func (s *Server) loadExtensions(exts []*ext.ExtensionInfo) error {
var lastError error
for _, extension := range exts {
if err := s.loadExtension(extension); err != nil {
lastError = err
}
}
return lastError
}
func (s *Server) loadExtension(extInfo *ext.ExtensionInfo) error {
if extInfo.ExtensionAPI != "v0" {
return fmt.Errorf("%s: unsupported extension API %s", extInfo.Name, extInfo.ExtensionAPI)
}
s.extensions = append(s.extensions, extInfo)
bitmapOps := extInfo.BitmapOps
bmOps, countOps, fieldOps, unknownOps := 0, 0, 0, 0
for i := range bitmapOps {
typ := bitmapOps[i].Func.BitmapOpType()
switch {
case typ.Input == ext.OpInputBitmap && typ.Output == ext.OpOutputCount:
countOps++
case typ.Input == ext.OpInputBitmap && typ.Output == ext.OpOutputBitmap:
bmOps++
case typ.Input == ext.OpInputNaryBSI && typ.Output == ext.OpOutputSignedBitmap:
fieldOps++
default:
unknownOps++
}
}
err := s.executor.registerOps(bitmapOps)
if err != nil {
s.logger.Printf("warning: extension registration failed: %v", err)
} else {
pql.RegisterPluginFuncs(bitmapOps)
}
return nil
}
// UpAndDown brings the server up minimally and shuts it down // UpAndDown brings the server up minimally and shuts it down
// again; basically, it exists for testing holder open and close. // again; basically, it exists for testing holder open and close.
func (s *Server) UpAndDown() error { func (s *Server) UpAndDown() error {

View file

@ -53,6 +53,9 @@ type Config struct {
// Bind is the host:port on which Pilosa will listen. // Bind is the host:port on which Pilosa will listen.
Bind string `toml:"bind"` Bind string `toml:"bind"`
// BindGRPC is the host:port on which Pilosa will bind for gRPC.
BindGRPC string `toml:"bind-grpc"`
// Advertise is the address advertised by the server to other nodes // Advertise is the address advertised by the server to other nodes
// in the cluster. It should be reachable by all other nodes and should // in the cluster. It should be reachable by all other nodes and should
// route to an interface that Bind is listening on. // route to an interface that Bind is listening on.
@ -161,6 +164,7 @@ func NewConfig() *Config {
c := &Config{ c := &Config{
DataDir: "~/.pilosa", DataDir: "~/.pilosa",
Bind: ":10101", Bind: ":10101",
BindGRPC: ":20101",
MaxWritesPerRequest: 5000, MaxWritesPerRequest: 5000,
// We default these Max File/Map counts very high. This is basically a // We default these Max File/Map counts very high. This is basically a
@ -233,6 +237,13 @@ func (cfg *Config) validateAddrs(ctx context.Context) error {
} }
cfg.Bind = schemeHostPortString(listenScheme, listenHost, listenPort) cfg.Bind = schemeHostPortString(listenScheme, listenHost, listenPort)
// Validate the gRPC listen address.
grpcListenScheme, grpcListenHost, grpcListenPort, err := validateListenAddr(ctx, cfg.BindGRPC)
if err != nil {
return errors.Wrap(err, "validating grpc listen address")
}
cfg.BindGRPC = schemeHostPortString(grpcListenScheme, grpcListenHost, grpcListenPort)
return nil return nil
} }

841
server/grpc.go Normal file
View file

@ -0,0 +1,841 @@
// Copyright 2017 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package server
import (
"context"
"crypto/tls"
"fmt"
"net"
"strings"
"github.com/pilosa/pilosa/v2"
"github.com/pilosa/pilosa/v2/logger"
pb "github.com/pilosa/pilosa/v2/proto"
"github.com/pkg/errors"
"google.golang.org/grpc"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/credentials"
"google.golang.org/grpc/reflection"
"google.golang.org/grpc/status"
)
// grpcHandler contains methods which handle the various gRPC requests.
type grpcHandler struct {
api *pilosa.API
logger logger.Logger
}
// errorToStatusError appends an appropriate grpc status code
// to the error (returning it as a status.Error). It is
// assumed that the input err is non-nil.
func errToStatusError(err error) error {
// Check error string.
switch errors.Cause(err) {
case pilosa.ErrIndexNotFound, pilosa.ErrFieldNotFound:
return status.Error(codes.NotFound, err.Error())
}
// Check error type.
switch errors.Cause(err).(type) {
case pilosa.NotFoundError:
return status.Error(codes.NotFound, err.Error())
}
return status.Error(codes.Unknown, err.Error())
}
// QueryPQL handles the PQL request and sends RowResponses to the stream.
func (h grpcHandler) QueryPQL(req *pb.QueryPQLRequest, stream pb.Pilosa_QueryPQLServer) error {
query := pilosa.QueryRequest{
Index: req.Index,
Query: req.Pql,
}
resp, err := h.api.Query(context.Background(), &query)
if err != nil {
return errToStatusError(err)
}
for row := range makeRows(resp, h.logger) {
err = stream.Send(row)
if err != nil {
return errToStatusError(err)
}
}
return nil
}
// fieldDataType returns a useful data type (string,
// uint64, bool, etc.) based on the Pilosa field type.
func fieldDataType(f *pilosa.Field) string {
switch f.Type() {
case "set", "mutex":
if f.Keys() {
return "[]string"
}
return "[]uint64"
case "int":
if f.Keys() {
return "string"
}
return "int64"
case "decimal":
return "float64"
case "bool":
return "bool"
case "time":
return "int64" // TODO: this is a placeholder
default:
panic(fmt.Sprintf("unimplemented fieldDataType: %s", f.Type()))
}
}
// Inspect handles the inspect request and sends an InspectResponse to the stream.
func (h grpcHandler) Inspect(req *pb.InspectRequest, stream pb.Pilosa_InspectServer) error {
const defaultLimit = 100000
index, err := h.api.Index(context.Background(), req.Index)
if err != nil {
return errToStatusError(err)
}
var fields []*pilosa.Field
for _, field := range index.Fields() {
// exclude internal fields (starting with "_")
if strings.HasPrefix(field.Name(), "_") {
continue
}
if len(req.FilterFields) > 0 {
for _, filter := range req.FilterFields {
if filter == field.Name() {
fields = append(fields, field)
break
}
}
} else {
fields = append(fields, field)
}
}
limit := req.Limit
if limit == 0 {
limit = defaultLimit
}
offset := req.Offset
if !index.Keys() {
ints, ok := req.Columns.Type.(*pb.IdsOrKeys_Ids)
if !ok {
return errors.New("invalid int columns")
}
ci := []*pb.ColumnInfo{
{Name: "_id", Datatype: "uint64"},
}
for _, field := range fields {
ci = append(ci, &pb.ColumnInfo{Name: field.Name(), Datatype: fieldDataType(field)})
}
// If Columns is empty, then get the _exists list (via All()),
// from the index and loop over that instead.
cols := ints.Ids.Vals
if len(cols) > 0 {
// Apply limit/offset to the provided columns.
if int(offset) >= len(cols) {
return nil
}
end := limit + offset
if int(end) > len(cols) {
end = uint64(len(cols))
}
cols = cols[offset:end]
} else {
// Prevent getting too many records by forcing a limit.
pql := fmt.Sprintf("All(limit=%d, offset=%d)", limit, offset)
query := pilosa.QueryRequest{
Index: req.Index,
Query: pql,
}
resp, err := h.api.Query(context.Background(), &query)
if err != nil {
return errors.Wrapf(err, "querying for all: %s", pql)
}
ids, ok := resp.Results[0].(*pilosa.Row)
if !ok {
return errors.Wrap(err, "getting results as a row")
}
limitedCols := ids.Columns()
if len(limitedCols) == 0 {
// If cols is still empty after the limit/offset, then
// return with no results.
return nil
}
cols = limitedCols
}
for _, col := range cols {
rowResp := &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: col}},
},
}
ci = nil // only include headers with the first row
for _, field := range fields {
// TODO: handle `time` fields
switch field.Type() {
case "set":
pql := fmt.Sprintf("Rows(%s, column=%d)", field.Name(), col)
query := pilosa.QueryRequest{
Index: req.Index,
Query: pql,
}
resp, err := h.api.Query(context.Background(), &query)
if err != nil {
return errors.Wrapf(err, "querying rows for set: %s", pql)
}
ids, ok := resp.Results[0].(pilosa.RowIdentifiers)
if !ok {
return errors.Wrap(err, "getting row identifiers")
}
if len(ids.Keys) > 0 {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringArrayVal{StringArrayVal: &pb.StringArray{Vals: ids.Keys}}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64ArrayVal{Uint64ArrayVal: &pb.Uint64Array{Vals: ids.Rows}}})
}
case "mutex":
pql := fmt.Sprintf("Rows(%s, column=%d)", field.Name(), col)
query := pilosa.QueryRequest{
Index: req.Index,
Query: pql,
}
resp, err := h.api.Query(context.Background(), &query)
if err != nil {
return errors.Wrap(err, "querying rows for mutex")
}
ids, ok := resp.Results[0].(pilosa.RowIdentifiers)
if !ok {
return errors.Wrap(err, "getting row identifiers")
}
if len(ids.Keys) == 1 {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringVal{StringVal: ids.Keys[0]}})
} else if len(ids.Rows) == 1 {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: ids.Rows[0]}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: nil})
}
case "int":
if field.Keys() {
value, exists, err := field.StringValue(col)
if err != nil {
return errors.Wrap(err, "getting string field value for column")
} else if exists {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringVal{StringVal: value}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: nil})
}
} else {
value, exists, err := field.Value(col)
if err != nil {
return errors.Wrap(err, "getting int field value for column")
} else if exists {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Int64Val{Int64Val: value}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: nil})
}
}
case "decimal":
value, exists, err := field.FloatValue(col)
if err != nil {
return errors.Wrap(err, "getting decimal field value for column")
} else if exists {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Float64Val{Float64Val: value}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: nil})
}
case "bool":
pql := fmt.Sprintf("Rows(%s, column=%d)", field.Name(), col)
query := pilosa.QueryRequest{
Index: req.Index,
Query: pql,
}
resp, err := h.api.Query(context.Background(), &query)
if err != nil {
return errors.Wrap(err, "querying rows for bool")
}
ids, ok := resp.Results[0].(pilosa.RowIdentifiers)
if !ok {
return errors.Wrap(err, "getting row identifiers")
}
if len(ids.Rows) == 1 {
var bval bool
if ids.Rows[0] == 1 {
bval = true
}
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_BoolVal{BoolVal: bval}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: nil})
}
case "time":
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: nil})
}
}
if err := stream.Send(rowResp); err != nil {
return errors.Wrap(err, "sending response to stream")
}
}
} else {
var cols []string
switch keys := req.Columns.Type.(type) {
case *pb.IdsOrKeys_Ids:
// The default behavior (in api/client/grpc.go) is to
// send an empty set of Ids even if the index supports
// keys, so in that case we just need to ignore it.
case *pb.IdsOrKeys_Keys:
cols = keys.Keys.Vals
default:
return errToStatusError(errors.New("invalid key columns"))
}
ci := []*pb.ColumnInfo{
{Name: "_id", Datatype: "string"},
}
for _, field := range fields {
ci = append(ci, &pb.ColumnInfo{Name: field.Name(), Datatype: fieldDataType(field)})
}
// If Columns is empty, then get the _exists list (via All()),
// from the index and loop over that instead.
if len(cols) > 0 {
// Apply limit/offset to the provided columns.
if int(offset) >= len(cols) {
return nil
}
end := limit + offset
if int(end) > len(cols) {
end = uint64(len(cols))
}
cols = cols[offset:end]
} else {
// Prevent getting too many records by forcing a limit.
pql := fmt.Sprintf("All(limit=%d, offset=%d)", limit, offset)
query := pilosa.QueryRequest{
Index: req.Index,
Query: pql,
}
resp, err := h.api.Query(context.Background(), &query)
if err != nil {
return errors.Wrapf(err, "querying for all: %s", pql)
}
ids, ok := resp.Results[0].(*pilosa.Row)
if !ok {
return errors.Wrap(err, "getting results as a row")
}
limitedCols := ids.Keys
if len(limitedCols) == 0 {
// If cols is still empty after the limit/offset, then
// return with no results.
return nil
}
cols = limitedCols
}
for _, col := range cols {
rowResp := &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringVal{StringVal: col}},
},
}
ci = nil // only include headers with the first row
for _, field := range fields {
// TODO: handle `time` fields
switch field.Type() {
case "set":
pql := fmt.Sprintf("Rows(%s, column=\"%s\")", field.Name(), col)
query := pilosa.QueryRequest{
Index: req.Index,
Query: pql,
}
resp, err := h.api.Query(context.Background(), &query)
if err != nil {
return errors.Wrap(err, "querying set rows(keys)")
}
ids, ok := resp.Results[0].(pilosa.RowIdentifiers)
if !ok {
return errors.Wrap(err, "getting row identifiers")
}
if len(ids.Keys) > 0 {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringArrayVal{StringArrayVal: &pb.StringArray{Vals: ids.Keys}}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64ArrayVal{Uint64ArrayVal: &pb.Uint64Array{Vals: ids.Rows}}})
}
case "mutex":
pql := fmt.Sprintf("Rows(%s, column=\"%s\")", field.Name(), col)
query := pilosa.QueryRequest{
Index: req.Index,
Query: pql,
}
resp, err := h.api.Query(context.Background(), &query)
if err != nil {
return errors.Wrap(err, "querying mutex rows(keys)")
}
ids, ok := resp.Results[0].(pilosa.RowIdentifiers)
if !ok {
return errors.Wrap(err, "getting row identifiers")
}
if len(ids.Keys) == 1 {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringVal{StringVal: ids.Keys[0]}})
} else if len(ids.Rows) == 1 {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: ids.Rows[0]}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: nil})
}
case "int":
// Translate column key.
id, err := index.TranslateStore().TranslateKey(col)
if err != nil {
return errors.Wrap(err, "translating column key")
}
if field.Keys() {
value, exists, err := field.StringValue(id)
if err != nil {
return errors.Wrap(err, "getting string field value for column")
} else if exists {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringVal{StringVal: value}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: nil})
}
} else {
value, exists, err := field.Value(id)
if err != nil {
return errors.Wrap(err, "getting int field value for column")
} else if exists {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Int64Val{Int64Val: value}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: nil})
}
}
case "decimal":
// Translate column key.
id, err := index.TranslateStore().TranslateKey(col)
if err != nil {
return errors.Wrap(err, "translating column key")
}
value, exists, err := field.FloatValue(id)
if err != nil {
return errors.Wrap(err, "getting decimal field value for column")
} else if exists {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Float64Val{Float64Val: value}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: nil})
}
case "bool":
pql := fmt.Sprintf("Rows(%s, column=\"%s\")", field.Name(), col)
query := pilosa.QueryRequest{
Index: req.Index,
Query: pql,
}
resp, err := h.api.Query(context.Background(), &query)
if err != nil {
return errors.Wrap(err, "querying bool rows(keys)")
}
ids, ok := resp.Results[0].(pilosa.RowIdentifiers)
if !ok {
return errors.Wrap(err, "getting row identifiers")
}
if len(ids.Rows) == 1 {
var bval bool
if ids.Rows[0] == 1 {
bval = true
}
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_BoolVal{BoolVal: bval}})
} else {
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: nil})
}
}
}
if err := stream.Send(rowResp); err != nil {
return errors.Wrap(err, "sending response to stream")
}
}
}
return nil
}
// I think ideally this would be plugged in the executor somewhere
// in order to get some concurrency benefit but we can
// start with the combined response
func makeRows(resp pilosa.QueryResponse, logger logger.Logger) chan *pb.RowResponse {
results := make(chan *pb.RowResponse)
go func() {
var breakLoop bool // Support the "break" inside the switch.
for _, result := range resp.Results {
if breakLoop {
break
}
switch r := result.(type) {
case *pilosa.Row:
if len(r.Keys) > 0 {
// Column keys
ci := []*pb.ColumnInfo{
{Name: "_id", Datatype: "string"},
}
for _, x := range r.Keys {
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringVal{StringVal: x}},
}}
ci = nil //only send on the first
}
} else {
// Column IDs
ci := []*pb.ColumnInfo{
{Name: "_id", Datatype: "uint64"},
}
for _, x := range r.Columns() {
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: x}},
}}
ci = nil //only send on the first
}
// The following will return roaring segments.
// This is commented out for now until we decide how we want to use this.
/*
// Roaring segments
ci := []*pb.ColumnInfo{
// TODO:
{Name: "shard", Datatype: "uint64"},
{Name: "segment", Datatype: "roaring"},
}
for _, x := range r.Segments() {
shard, b := x.Raw()
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_IntVal{int64(shard)}},
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_BlobVal{b}},
}}
ci = nil //only send on the first
}
*/
}
case pilosa.PairField:
if r.Pair.Key != "" {
results <- &pb.RowResponse{
Headers: []*pb.ColumnInfo{
{Name: r.Field, Datatype: "string"},
{Name: "count", Datatype: "uint64"},
},
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringVal{StringVal: r.Pair.Key}},
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: r.Pair.Count}},
},
}
} else {
results <- &pb.RowResponse{
Headers: []*pb.ColumnInfo{
{Name: r.Field, Datatype: "uint64"},
{Name: "count", Datatype: "uint64"},
},
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: r.Pair.ID}},
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: r.Pair.Count}},
},
}
}
case *pilosa.PairsField:
// Determine if the ID has string keys.
var stringKeys bool
if len(r.Pairs) > 0 {
if r.Pairs[0].Key != "" {
stringKeys = true
}
}
dtype := "uint64"
if stringKeys {
dtype = "string"
}
ci := []*pb.ColumnInfo{
{Name: r.Field, Datatype: dtype},
{Name: "count", Datatype: "uint64"},
}
for _, pair := range r.Pairs {
if stringKeys {
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringVal{StringVal: pair.Key}},
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: uint64(pair.Count)}},
},
}
} else {
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: uint64(pair.ID)}},
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: uint64(pair.Count)}},
},
}
}
ci = nil //only send on the first
}
case []pilosa.GroupCount:
for i, gc := range r {
var ci []*pb.ColumnInfo
if i == 0 {
for _, fieldRow := range gc.Group {
if fieldRow.RowKey != "" {
ci = append(ci, &pb.ColumnInfo{Name: fieldRow.Field, Datatype: "string"})
} else {
ci = append(ci, &pb.ColumnInfo{Name: fieldRow.Field, Datatype: "uint64"})
}
}
ci = append(ci, &pb.ColumnInfo{Name: "count", Datatype: "uint64"})
ci = append(ci, &pb.ColumnInfo{Name: "sum", Datatype: "int64"})
}
rowResp := &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{},
}
for _, fieldRow := range gc.Group {
if fieldRow.RowKey != "" {
rowResp.Columns = append(rowResp.Columns, &pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringVal{StringVal: fieldRow.RowKey}})
} else {
rowResp.Columns = append(rowResp.Columns, &pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: uint64(fieldRow.RowID)}})
}
}
rowResp.Columns = append(rowResp.Columns,
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: gc.Count}},
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Int64Val{Int64Val: gc.Sum}},
)
results <- rowResp
}
case pilosa.RowIdentifiers:
if len(r.Keys) > 0 {
ci := []*pb.ColumnInfo{{Name: r.Field(), Datatype: "string"}}
for _, key := range r.Keys {
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_StringVal{StringVal: key}},
}}
ci = nil
}
} else {
ci := []*pb.ColumnInfo{{Name: r.Field(), Datatype: "uint64"}}
for _, id := range r.Rows {
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: uint64(id)}},
}}
ci = nil
}
}
case uint64:
ci := []*pb.ColumnInfo{{Name: "count", Datatype: "uint64"}}
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Uint64Val{Uint64Val: uint64(r)}},
}}
case bool:
ci := []*pb.ColumnInfo{{Name: "result", Datatype: "bool"}}
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_BoolVal{BoolVal: r}},
}}
case pilosa.ValCount:
ci := []*pb.ColumnInfo{
{Name: "value", Datatype: "int64"},
{Name: "count", Datatype: "int64"},
}
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Int64Val{Int64Val: r.Val}},
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Int64Val{Int64Val: r.Count}},
}}
case pilosa.SignedRow:
// TODO: address the overflow issue with values outside the int64 range
ci := []*pb.ColumnInfo{{Name: r.Field(), Datatype: "int64"}}
negs := r.Neg.Columns()
for i := len(negs) - 1; i >= 0; i-- {
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Int64Val{Int64Val: -1 * int64(negs[i])}},
}}
ci = nil
}
for _, id := range r.Pos.Columns() {
results <- &pb.RowResponse{
Headers: ci,
Columns: []*pb.ColumnResponse{
&pb.ColumnResponse{ColumnVal: &pb.ColumnResponse_Int64Val{Int64Val: int64(id)}},
}}
ci = nil
}
default:
logger.Printf("unhandled %T\n", r)
breakLoop = true
}
}
close(results)
}()
return results
}
type grpcServer struct {
api *pilosa.API
grpcServer *grpc.Server
hostPort string
logger logger.Logger
}
type grpcServerOption func(s *grpcServer) error
func OptGRPCServerAPI(api *pilosa.API) grpcServerOption {
return func(s *grpcServer) error {
s.api = api
return nil
}
}
func OptGRPCServerURI(uri *pilosa.URI) grpcServerOption {
hostport := fmt.Sprintf("%s:%d", uri.Host, uri.Port)
return func(s *grpcServer) error {
s.hostPort = hostport
return nil
}
}
func OptGRPCServerLogger(logger logger.Logger) grpcServerOption {
return func(s *grpcServer) error {
s.logger = logger
return nil
}
}
func (s *grpcServer) Serve(tlsConfig *tls.Config) error {
// create listener
lis, err := net.Listen("tcp", s.hostPort)
if err != nil {
return errors.Wrap(err, "creating listener")
}
s.logger.Printf("enabled grpc listening on %s", lis.Addr())
opts := make([]grpc.ServerOption, 0)
if tlsConfig != nil {
creds := credentials.NewTLS(tlsConfig)
opts = append(opts, grpc.Creds(creds))
}
// create grpc server
s.grpcServer = grpc.NewServer(opts...)
pb.RegisterPilosaServer(s.grpcServer, grpcHandler{api: s.api, logger: s.logger})
// register the server so its services are available to grpc_cli and others
reflection.Register(s.grpcServer)
// and start...
if err := s.grpcServer.Serve(lis); err != nil {
return errors.Wrap(err, "starting grpc server")
}
return nil
}
func NewGRPCServer(opts ...grpcServerOption) (*grpcServer, error) {
server := &grpcServer{
logger: logger.NopLogger,
}
for _, opt := range opts {
err := opt(server)
if err != nil {
return nil, errors.Wrap(err, "applying option")
}
}
return server, nil
}

View file

@ -0,0 +1,278 @@
// Copyright 2017 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package server
import (
"testing"
"github.com/pilosa/pilosa/v2"
"github.com/pilosa/pilosa/v2/logger"
)
func TestGRPC(t *testing.T) {
t.Run("makeRows", func(t *testing.T) {
type expHeader struct {
name string
dataType string
}
type expColumn interface{}
tests := []struct {
result interface{}
expHeaders []expHeader
expColumns [][]expColumn
}{
// Row (uint64)
{
pilosa.NewRow(10, 11, 12),
[]expHeader{
{"_id", "uint64"},
},
[][]expColumn{
{uint64(10)},
{uint64(11)},
{uint64(12)},
},
},
// Row (string)
{
&pilosa.Row{Keys: []string{"ten", "eleven", "twelve"}},
[]expHeader{
{"_id", "string"},
},
[][]expColumn{
{"ten"},
{"eleven"},
{"twelve"},
},
},
// Pair (uint64)
{
pilosa.Pair{ID: 10, Count: 123},
[]expHeader{
{"_id", "uint64"},
{"count", "uint64"},
},
[][]expColumn{
{uint64(10), uint64(123)},
},
},
// Pair (string)
{
pilosa.Pair{Key: "ten", Count: 123},
[]expHeader{
{"_id", "string"},
{"count", "uint64"},
},
[][]expColumn{
{string("ten"), uint64(123)},
},
},
// []Pair (uint64)
{
[]pilosa.Pair{
{ID: 10, Count: 123},
{ID: 11, Count: 456},
},
[]expHeader{
{"_id", "uint64"},
{"count", "uint64"},
},
[][]expColumn{
{uint64(10), uint64(123)},
{uint64(11), uint64(456)},
},
},
// []Pair (string)
{
[]pilosa.Pair{
{Key: "ten", Count: 123},
{Key: "eleven", Count: 456},
},
[]expHeader{
{"_id", "string"},
{"count", "uint64"},
},
[][]expColumn{
{"ten", uint64(123)},
{"eleven", uint64(456)},
},
},
// []GroupCount (uint64)
{
[]pilosa.GroupCount{
pilosa.GroupCount{
Group: []pilosa.FieldRow{
{Field: "a", RowID: 10},
{Field: "b", RowID: 11},
},
Count: 123,
},
pilosa.GroupCount{
Group: []pilosa.FieldRow{
{Field: "a", RowID: 10},
{Field: "b", RowID: 12},
},
Count: 456,
},
},
[]expHeader{
{"a", "uint64"},
{"b", "uint64"},
{"count", "uint64"},
{"sum", "int64"},
},
[][]expColumn{
{uint64(10), uint64(11), uint64(123), int64(0)},
{uint64(10), uint64(12), uint64(456), int64(0)},
},
},
// []GroupCount (string)
{
[]pilosa.GroupCount{
pilosa.GroupCount{
Group: []pilosa.FieldRow{
{Field: "a", RowKey: "ten"},
{Field: "b", RowKey: "eleven"},
},
Count: 123,
},
{
Group: []pilosa.FieldRow{
{Field: "a", RowKey: "ten"},
{Field: "b", RowKey: "twelve"},
},
Count: 456,
},
},
[]expHeader{
{"a", "string"},
{"b", "string"},
{"count", "uint64"},
{"sum", "int64"},
},
[][]expColumn{
{"ten", "eleven", uint64(123), int64(0)},
{"ten", "twelve", uint64(456), int64(0)},
},
},
// RowIdentifiers (uint64)
{
pilosa.RowIdentifiers{
Rows: []uint64{10, 11, 12},
},
[]expHeader{
{"", "uint64"}, // This is blank because we don't expose RowIdentifiers.field, so we have no way to set it for tests.
},
[][]expColumn{
{uint64(10)},
{uint64(11)},
{uint64(12)},
},
},
// RowIdentifiers (string)
{
pilosa.RowIdentifiers{
Keys: []string{"ten", "eleven", "twelve"},
},
[]expHeader{
{"", "string"}, // This is blank because we don't expose RowIdentifiers.field, so we have no way to set it for tests.
},
[][]expColumn{
{"ten"},
{"eleven"},
{"twelve"},
},
},
// uint64
{
uint64(123),
[]expHeader{
{"count", "uint64"},
},
[][]expColumn{
{uint64(123)},
},
},
// bool
{
true,
[]expHeader{
{"result", "bool"},
},
[][]expColumn{
{true},
},
},
}
logger := logger.NopLogger
for ti, test := range tests {
results := make([]interface{}, 0)
results = append(results, test.result)
qr := pilosa.QueryResponse{}
qr.Results = results
ch := makeRows(qr, logger)
cnt := 0
for row := range ch {
// Ensure headers match (on the first row).
if cnt == 0 {
for i, header := range row.GetHeaders() {
if header.Name != test.expHeaders[i].name {
t.Fatalf("test %d expected header name: %s, but got: %s", ti, test.expHeaders[i].name, header.Name)
}
if header.Datatype != test.expHeaders[i].dataType {
t.Fatalf("test %d expected header data type: %s, but got: %s", ti, test.expHeaders[i].dataType, header.Datatype)
}
}
}
// Ensure column data matches.
for i, column := range row.GetColumns() {
switch v := test.expColumns[cnt][i].(type) {
case string:
val := column.GetStringVal()
if val != v {
t.Fatalf("test %d expected column val: %v, but got: %v", ti, v, val)
}
case uint64:
val := column.GetUint64Val()
if val != v {
t.Fatalf("test %d expected column val: %v, but got: %v", ti, v, val)
}
case bool:
val := column.GetBoolVal()
if val != v {
t.Fatalf("test %d expected column val: %v, but got: %v", ti, v, val)
}
case int64:
val := column.GetInt64Val()
if val != v {
t.Fatalf("test %d expected column val: %v but got: %v", ti, v, val)
}
default:
t.Fatalf("test %d has unhandled data type: %T", ti, v)
}
}
cnt++
}
}
})
}

View file

@ -224,7 +224,12 @@ func TestHandler_Endpoints(t *testing.T) {
if err != nil { if err != nil {
t.Fatalf("querying: %v", err) t.Fatalf("querying: %v", err)
} }
if !reflect.DeepEqual(resp.Results[0], []pilosa.Pair{{Count: 12, ID: 0}}) { if !reflect.DeepEqual(resp.Results[0], &pilosa.PairsField{
Pairs: []pilosa.Pair{
{Count: 12, ID: 0},
},
Field: "f1",
}) {
t.Fatalf("Unexpected result %v", resp.Results[0]) t.Fatalf("Unexpected result %v", resp.Results[0])
} }
@ -504,8 +509,8 @@ func TestHandler_Endpoints(t *testing.T) {
var resp pilosa.QueryResponse var resp pilosa.QueryResponse
if err := cmd.API.Serializer.Unmarshal(w.Body.Bytes(), &resp); err != nil { if err := cmd.API.Serializer.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatal(err) t.Fatal(err)
} else if a := resp.Results[0].([]pilosa.Pair); len(a) != 2 { } else if a := resp.Results[0].(*pilosa.PairsField); len(a.Pairs) != 2 {
t.Fatalf("unexpected pair length: %d", len(a)) t.Fatalf("unexpected pair length: %d", len(a.Pairs))
} }
}) })

View file

@ -80,9 +80,11 @@ type Command struct {
logger loggerLogger logger loggerLogger
Handler pilosa.Handler Handler pilosa.Handler
grpcServer *grpcServer
API *pilosa.API API *pilosa.API
ln net.Listener ln net.Listener
listenURI *pilosa.URI listenURI *pilosa.URI
tlsConfig *tls.Config
closeTimeout time.Duration closeTimeout time.Duration
serverOptions []pilosa.ServerOption serverOptions []pilosa.ServerOption
@ -163,6 +165,11 @@ func (m *Command) Start() (err error) {
} }
m.logger.Printf("listening as %s\n", m.listenURI) m.logger.Printf("listening as %s\n", m.listenURI)
go func() {
if err := m.grpcServer.Serve(m.tlsConfig); err != nil {
m.logger.Printf("grpc server error: %v", err)
}
}()
close(m.Started) close(m.Started)
return nil return nil
@ -251,10 +258,14 @@ func (m *Command) SetupServer() error {
return errors.Wrap(err, "processing bind address") return errors.Wrap(err, "processing bind address")
} }
grpcURI, err := pilosa.NewURIFromAddress(m.Config.BindGRPC)
if err != nil {
return errors.Wrap(err, "processing bind grpc address")
}
// Setup TLS // Setup TLS
var TLSConfig *tls.Config
if uri.Scheme == "https" { if uri.Scheme == "https" {
TLSConfig, err = GetTLSConfig(&m.Config.TLS, m.logger.Logger()) m.tlsConfig, err = GetTLSConfig(&m.Config.TLS, m.logger.Logger())
if err != nil { if err != nil {
return errors.Wrap(err, "get tls config") return errors.Wrap(err, "get tls config")
} }
@ -270,7 +281,7 @@ func (m *Command) SetupServer() error {
return errors.Wrap(err, "new stats client") return errors.Wrap(err, "new stats client")
} }
m.ln, err = getListener(*uri, TLSConfig) m.ln, err = getListener(*uri, m.tlsConfig)
if err != nil { if err != nil {
return errors.Wrap(err, "getting listener") return errors.Wrap(err, "getting listener")
} }
@ -283,7 +294,7 @@ func (m *Command) SetupServer() error {
// Save listenURI for later reference. // Save listenURI for later reference.
m.listenURI = uri m.listenURI = uri
c := http.GetHTTPClient(TLSConfig) c := http.GetHTTPClient(m.tlsConfig)
// Get advertise address as uri. // Get advertise address as uri.
advertiseURI, err := pilosa.AddressWithDefaults(m.Config.Advertise) advertiseURI, err := pilosa.AddressWithDefaults(m.Config.Advertise)
@ -351,7 +362,16 @@ func (m *Command) SetupServer() error {
http.OptHandlerListener(m.ln), http.OptHandlerListener(m.ln),
http.OptHandlerCloseTimeout(m.closeTimeout), http.OptHandlerCloseTimeout(m.closeTimeout),
) )
if err != nil {
return errors.Wrap(err, "new handler") return errors.Wrap(err, "new handler")
}
m.grpcServer, err = NewGRPCServer(
OptGRPCServerAPI(m.API),
OptGRPCServerURI(grpcURI),
OptGRPCServerLogger(m.logger),
)
return errors.Wrap(err, "new grpc server")
} }
// setupNetworking sets up internode communication based on the configuration. // setupNetworking sets up internode communication based on the configuration.

406
snapshotqueue.go Normal file
View file

@ -0,0 +1,406 @@
// Copyright 2019 Pilosa Corp.
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package pilosa
import (
"fmt"
"os"
"sync"
"sync/atomic"
"time"
"github.com/pilosa/pilosa/v2/logger"
"github.com/pkg/errors"
)
// snapshotQueue is a thing which can handle enqueuing snapshots. A snapshot
// queue distinguishes between high-priority requests, which get satisfied
// by the next available worker, and regular requests, which get enqueued
// if there's space in the queue, and otherwise dropped. There's also a
// separate background task to scan a holder for fragments which may need
// snapshots, but which is processed only when the queue is empty, and only
// slowly. "Await" awaits an existing snapshot if one is already enqueued.
// "Immediate" tries to do one right away. (If one's already enqueued, this
// can leave it in the queue, which will ignore anything that shows up with
// the request flag cleared.)
//
// Await, Enqueue, and Immediate should be called only with the fragment lock
// held.
//
// ScanHolder spawns a new goroutine. You don't need to use `go` on it.
type snapshotQueue interface {
Immediate(*fragment) error
Enqueue(*fragment)
Await(*fragment) error
ScanHolder(*Holder)
Stop()
}
// queuelessSnapshotQueue isn't a snapshot queue, but it satisfies the
// interface.
type queuelessSnapshotQueue struct{}
func (q *queuelessSnapshotQueue) Enqueue(f *fragment) {
_ = f.snapshot()
}
func (q *queuelessSnapshotQueue) Await(f *fragment) error {
return nil
}
func (q *queuelessSnapshotQueue) Immediate(f *fragment) error {
return f.snapshot()
}
func (q *queuelessSnapshotQueue) ScanHolder(h *Holder) {
}
func (q *queuelessSnapshotQueue) Stop() {
}
// defaultSnapshotQueue is the fallback to use if none is available,
// and currently uses queueless -- it runs all snapshots immediately.
var defaultSnapshotQueue *queuelessSnapshotQueue
// newSnapshotQueue makes a new snapshot queue, of depth N, with
// w worker threads.
func newSnapshotQueue(n int, w int, l logger.Logger) snapshotQueue {
sq := prioritySnapshotQueue{normal: make(chan snapshotRequest, n), urgent: make(chan snapshotRequest), background: make(chan snapshotRequest), done: make(chan struct{}), logger: l}
if sq.logger == nil {
sq.logger = logger.NewStandardLogger(os.Stderr)
}
sq.spawnWorkers(w)
return &sq
}
type snapshotRequest struct {
frag *fragment
when time.Time
}
// prioritySnapshotQueue gives preference to "immediate" requests, and
// dispreference to "background" requests from ScanHolder. It timestamps
// requests, so it can discard a request if the most recent snapshot is
// newer than the request. The snapshotPending flag in the fragment is
// used to track that a given fragment thinks it has been successfully
// enqueued. Background requests are not considered enqueued, since
// they'll never get processed if there's anything else. In normal workloads,
// immediate/urgent snapshots should be rare, but we'll happily drop
// most requests on the floor; the scanner should pick them up once things
// are quiet.
type prioritySnapshotQueue struct {
logger logger.Logger
urgent chan snapshotRequest
normal chan snapshotRequest
background chan snapshotRequest
done chan struct{}
mu sync.RWMutex
scanWG, workerWG sync.WaitGroup
stats struct {
enqueued uint64
skipped uint64
}
}
func (sq *prioritySnapshotQueue) spawnWorkers(w int) {
sq.mu.Lock()
defer sq.mu.Unlock()
if sq.done == nil {
sq.logger.Printf("prioritySnapshotQueue worker: no done channel, already done?")
return
}
sq.workerWG.Add(w)
for i := 0; i < w; i++ {
go sq.worker(sq.urgent, sq.normal, sq.background, sq.done)
}
}
func (sq *prioritySnapshotQueue) worker(urgent, normal, background chan snapshotRequest, done chan struct{}) {
// We don't want a race condition on these. If they're non-nil when
// we get them, they should get closed at some point. If done is
// already nil, we shouldn't do anything.
defer sq.workerWG.Done()
ok := true
var req snapshotRequest
for ok {
req.frag = nil
select {
case req, ok = <-urgent:
default:
select {
case req, ok = <-urgent:
case req, ok = <-normal:
default:
select {
case req, ok = <-urgent:
case req, ok = <-normal:
case req, ok = <-background:
case _, ok = <-done:
}
}
}
if req.frag != nil {
sq.process(req)
}
}
}
// process actually runs a fragment. it will do this if either the fragment
// has a pending snapshot, or the force flag is set.
func (sq *prioritySnapshotQueue) process(req snapshotRequest) {
f := req.frag
f.mu.Lock()
defer f.mu.Unlock()
if f.snapshotStamp.Before(req.when) {
f.snapshotErr = f.snapshot()
if f.snapshotErr != nil {
fmt.Printf("snapshot error: %v\n", f.snapshotErr)
sq.logger.Printf("snapshot error: %v", f.snapshotErr)
}
f.snapshotPending = false
f.snapshotCond.Broadcast()
}
}
// Stop shuts down the snapshot queue. It first marks it as done, causing
// the background scanner(s), if any, to shut down, then waits for them, then
// closes and nils the queues. The background scanner has to get stopped
// because otherwise it might try to write to those closed queues.
func (sq *prioritySnapshotQueue) Stop() {
sq.mu.Lock()
defer sq.mu.Unlock()
close(sq.done)
// scanners need to be done before we close the other channels.
sq.scanWG.Wait()
sq.done = nil
close(sq.normal)
sq.normal = nil
close(sq.urgent)
sq.urgent = nil
close(sq.background)
sq.background = nil
sq.logger.Printf("snapshot queue: enqueued %d, skipped %d\n", sq.stats.enqueued, sq.stats.skipped)
}
// Enqueue tries to add a fragment to the queue, if the fragment is not already
// enqueued. You should hold a lock on the fragment when calling this.
func (sq *prioritySnapshotQueue) Enqueue(f *fragment) {
if f.snapshotPending {
return
}
sq.mu.RLock()
defer sq.mu.RUnlock()
if sq.normal == nil {
sq.logger.Printf("requested snapshot after snapshot queue was closed")
return
}
// we have to set this before enqueing, because it's
// otherwise possible that we're at the head of the queue,
// and the recipient gets the fragment before we execute the
// line after the send.
f.snapshotPending = true
// try to enqueue snapshot
select {
case sq.normal <- snapshotRequest{frag: f, when: time.Now()}:
atomic.AddUint64(&sq.stats.enqueued, 1)
return
default:
atomic.AddUint64(&sq.stats.skipped, 1)
f.snapshotPending = false
return
}
}
// Await returns when f is not pending a snapshot. Call with the fragment lock
// held. Await waits on a condition variable inside f, associated with the
// fragment's lock, so this does not conflict with the lock being used for
// snapshots.
func (sq *prioritySnapshotQueue) Await(f *fragment) (err error) {
for f.snapshotPending {
f.snapshotCond.Wait()
}
err, f.snapshotErr = f.snapshotErr, nil
return err
}
// Immediate forces an immediate snapshot of the given fragment. Call with
// the fragment locked. If the queue is already closing, the fragment does
// not get snapshotted.
func (sq *prioritySnapshotQueue) Immediate(f *fragment) error {
sq.mu.RLock()
// no deferred unlock, because we want to unlock this before calling Await.
// Not because that needs this lock, but because once we're that far, we
// *don't* need this lock anymore so someone else should have it.
if sq.urgent == nil {
sq.mu.RUnlock()
sq.logger.Printf("requested immediate snapshot after snapshot queue was closed")
return errors.New("requested immediate snapshot after snapshot queue was closed")
}
f.snapshotPending = true
req := snapshotRequest{frag: f, when: time.Now()}
// if the fragment was already in the work queue, it's *possible*
// that the only available worker just picked it off the queue, and
// is now waiting on getting the fragment's lock, so it can run
// a snapshot. So we let go of the lock on the fragment, send the
// request, then request the fragment lock again, because Await will
// be sleeping on the condition variable associated with the lock,
// which means it needs to hold the lock so it can let it go during
// the wait... No, really, this made sense.
f.mu.Unlock()
sq.urgent <- req
sq.mu.RUnlock()
f.mu.Lock()
return sq.Await(f)
}
// needsSnapshot determines whether a fragment probably wants snapshotting.
// Specifically, it looks for fragments not already marked to receive
// snapshots, but which have a high enough opN to justify a snapshot. This
// is only used from the background scan.
func (sq *prioritySnapshotQueue) needsSnapshot(f *fragment) bool {
if f == nil {
return false
}
f.mu.Lock()
defer f.mu.Unlock()
if f.snapshotPending {
return false
}
if f.opN > f.MaxOpN {
return true
}
return false
}
// ScanHolder spawns a goroutine which iterates through the holder's
// indexes/fields/views/fragments, looking for fragments which have OpN
// high enough to justify a snapshot but don't seem to have one pending.
// It then dumps these in the low priority background queue.
func (sq *prioritySnapshotQueue) ScanHolder(h *Holder) {
sq.mu.Lock()
sq.scanWG.Add(1)
go sq.scanHolderWorker(h, sq.background, sq.done)
sq.mu.Unlock()
}
// scanHolderWorker is a background task that scans a holder looking for
// fragments which need snapshots taken. It's the cleanup task for snapshots
// that would have been requested by Enqueue, but the queue was full.
func (sq *prioritySnapshotQueue) scanHolderWorker(h *Holder, background chan snapshotRequest, done chan struct{}) {
defer sq.scanWG.Done()
var indexNames, fieldNames, viewNames []string
var fragNums []uint64
for {
// To avoid abusing things, cap activity rate; every time we finish
// the holder, or every couple hundred fragments considered, we
// pause for a bit.
counter := 0
hits := 0
h.mu.Lock()
indexNames = indexNames[:0]
for indexName := range h.indexes {
indexNames = append(indexNames, indexName)
}
h.mu.Unlock()
for _, indexName := range indexNames {
h.mu.Lock()
index := h.indexes[indexName]
h.mu.Unlock()
if index == nil {
continue
}
fieldNames = fieldNames[:0]
index.mu.Lock()
for fieldName := range index.fields {
fieldNames = append(fieldNames, fieldName)
}
index.mu.Unlock()
for _, fieldName := range fieldNames {
index.mu.Lock()
field := index.fields[fieldName]
index.mu.Unlock()
if field == nil {
continue
}
viewNames = viewNames[:0]
field.mu.Lock()
for viewName := range field.viewMap {
viewNames = append(viewNames, viewName)
}
field.mu.Unlock()
for _, viewName := range viewNames {
field.mu.Lock()
view := field.viewMap[viewName]
field.mu.Unlock()
if view == nil {
continue
}
fragNums := fragNums[:0]
view.mu.Lock()
for fragNum := range view.fragments {
fragNums = append(fragNums, fragNum)
}
view.mu.Unlock()
for _, fragNum := range fragNums {
view.mu.Lock()
frag := view.fragments[fragNum]
view.mu.Unlock()
if sq.needsSnapshot(frag) {
hits++
select {
case background <- snapshotRequest{frag: frag, when: time.Now()}:
sq.logger.Debugf("found fragment needing snapshot: %s\n", frag.path)
case <-done:
return
}
} else {
// Count fragments examined *without* finding anything that
// needed a snapshot. When we find things that need snapshots,
// the time it takes the workers to respond to us is enough
// of a delay to keep us from eating every CPU. So, if a lot
// of things need snapshots, and the workers aren't doing
// anything else, ScanHolder will mostly keep them saturated.
// If they're busy, we'll block forever in the write to the
// background queue. If there's nothing that needs snapshots,
// we pause frequently for a second or so at a time.
counter++
if counter == 100 {
select {
case <-time.After(1 * time.Second):
case <-done:
return
}
counter = 0
}
}
}
}
}
}
if hits > 0 {
sq.logger.Printf("background scan: %d fragments needed snapshots\n", hits)
hits = 0
} else {
sq.logger.Debugf("background scan: no fragments needed snapshots, waiting\n")
// No reason to be active if we're not finding anything.
select {
case <-time.After(60 * time.Second):
case <-done:
return
}
}
}
}

View file

@ -25,7 +25,7 @@ import (
) )
// Expvar global expvar map. // Expvar global expvar map.
var Expvar = expvar.NewMap("index") var Expvar *expvar.Map
// StatsClient represents a client to a stats server. // StatsClient represents a client to a stats server.
type StatsClient interface { type StatsClient interface {
@ -90,6 +90,9 @@ type expvarStatsClient struct {
// NewExpvarStatsClient returns a new instance of ExpvarStatsClient. // NewExpvarStatsClient returns a new instance of ExpvarStatsClient.
// This client points at the root of the expvar index map. // This client points at the root of the expvar index map.
func NewExpvarStatsClient() *expvarStatsClient { func NewExpvarStatsClient() *expvarStatsClient {
if Expvar == nil {
Expvar = expvar.NewMap("index")
}
return &expvarStatsClient{ return &expvarStatsClient{
m: Expvar, m: Expvar,
} }

View file

@ -18,12 +18,14 @@ import (
"bytes" "bytes"
"fmt" "fmt"
"io/ioutil" "io/ioutil"
"sync"
) )
// bufferLogger represents a test Logger that holds log messages // bufferLogger represents a test Logger that holds log messages
// in a buffer for review. // in a buffer for review.
type bufferLogger struct { type bufferLogger struct {
buf *bytes.Buffer buf *bytes.Buffer
mu sync.Mutex
} }
// NewBufferLogger returns a new instance of BufferLogger. // NewBufferLogger returns a new instance of BufferLogger.
@ -34,6 +36,8 @@ func NewBufferLogger() *bufferLogger {
} }
func (b *bufferLogger) Printf(format string, v ...interface{}) { func (b *bufferLogger) Printf(format string, v ...interface{}) {
b.mu.Lock()
defer b.mu.Unlock()
s := fmt.Sprintf(format, v...) s := fmt.Sprintf(format, v...)
_, err := b.buf.WriteString(s) _, err := b.buf.WriteString(s)
if err != nil { if err != nil {
@ -44,5 +48,7 @@ func (b *bufferLogger) Printf(format string, v ...interface{}) {
func (b *bufferLogger) Debugf(format string, v ...interface{}) {} func (b *bufferLogger) Debugf(format string, v ...interface{}) {}
func (b *bufferLogger) ReadAll() ([]byte, error) { func (b *bufferLogger) ReadAll() ([]byte, error) {
b.mu.Lock()
defer b.mu.Unlock()
return ioutil.ReadAll(b.buf) return ioutil.ReadAll(b.buf)
} }

View file

@ -73,6 +73,9 @@ func newCommand(opts ...server.CommandOption) *Command {
if m.Config.Bind == defaultConf.Bind { if m.Config.Bind == defaultConf.Bind {
m.Config.Bind = "http://localhost:0" m.Config.Bind = "http://localhost:0"
} }
if m.Config.BindGRPC == defaultConf.BindGRPC {
m.Config.BindGRPC = "http://localhost:0"
}
m.Config.Translation.MapSize = 140000 m.Config.Translation.MapSize = 140000
m.Config.WorkerPoolSize = 2 m.Config.WorkerPoolSize = 2

View file

@ -17,15 +17,51 @@ package tracing
import ( import (
"context" "context"
"net/http" "net/http"
"time"
) )
// GlobalTracer is a single, global instance of Tracer. // GlobalTracer is a single, global instance of Tracer.
var GlobalTracer Tracer = NopTracer() var GlobalTracer Tracer = NopTracer()
// StartSpanFromContext returnus a new child span and context from a given // StartSpanFromContext returns a new child span and context from a given
// context using the global tracer. // context using the global tracer.
func StartSpanFromContext(ctx context.Context, operationName string) (Span, context.Context) { func StartSpanFromContext(ctx context.Context, operationName string) (Span, context.Context) {
return startMaybeProfiledSpanFromContext(ctx, operationName, false)
}
// StartProfiledSpanFromContext returns a new child span and context from a given
// context using the global tracer.
func StartProfiledSpanFromContext(ctx context.Context, operationName string) (ProfiledSpan, context.Context) {
span, ctx := startMaybeProfiledSpanFromContext(ctx, operationName, true)
return span.(ProfiledSpan), ctx
}
// startMaybeProfiledSpanFromContext figures out whether it needs to make a profiling span.
func startMaybeProfiledSpanFromContext(ctx context.Context, operationName string, startProfiling bool) (Span, context.Context) {
var parent ProfiledSpan
makeProfile := startProfiling
// Parent context may or may not be profiled.
// If it is, we need to make a sub-profile. If it isn't, we need to make a new profile if
// startProfiling is set.
span, ok := spanFromContext(ctx)
if ok {
if parent, ok = span.(ProfiledSpan); ok {
makeProfile = true
}
}
// if we aren't making a profile, this is easy:
if !makeProfile {
return GlobalTracer.StartSpanFromContext(ctx, operationName) return GlobalTracer.StartSpanFromContext(ctx, operationName)
}
newProf := &Profile{Name: operationName, Begin: time.Now(), KV: make(map[string]interface{})}
if parent != nil {
parent.AddChild(newProf)
}
inner, ctx := GlobalTracer.StartSpanFromContext(ctx, operationName)
newProf.inner = inner
// insert ourselves in the context
ctx = context.WithValue(ctx, arbitraryContextKey, newProf)
return newProf, ctx
} }
// Tracer implements a generic distributed tracing interface. // Tracer implements a generic distributed tracing interface.
@ -49,6 +85,49 @@ type Span interface {
LogKV(alternatingKeyValues ...interface{}) LogKV(alternatingKeyValues ...interface{})
} }
// ProfiledSpan represents a span which profiles itself and its children.
type ProfiledSpan interface {
Span
Dump() interface{} // suitable for marshaling
AddChild(ProfiledSpan)
}
// Profile represents the profiling data for a span. It also handles
// the bookkeeping for the underlying span, but this is unexported so
// it doesn't get unmarshaled.
type Profile struct {
inner Span
Name string
Begin, End time.Time `json:"-"`
Duration time.Duration
Children []ProfiledSpan `json:",omitempty"`
KV map[string]interface{} `json:",omitempty"`
}
func (p *Profile) Finish() {
p.inner.Finish()
p.End = time.Now()
p.Duration = p.End.Sub(p.Begin)
}
func (p *Profile) LogKV(alternatingKeyValues ...interface{}) {
for i := 0; i < len(alternatingKeyValues)-1; i += 2 {
if s, ok := alternatingKeyValues[i].(string); ok {
p.KV[s] = alternatingKeyValues[i+1]
}
}
p.inner.LogKV(alternatingKeyValues...)
}
// returns something that json could probably marshal.
func (p *Profile) Dump() interface{} {
return p
}
func (p *Profile) AddChild(child ProfiledSpan) {
p.Children = append(p.Children, child)
}
// NopTracer returns a tracer that doesn't do anything. // NopTracer returns a tracer that doesn't do anything.
func NopTracer() Tracer { func NopTracer() Tracer {
return &nopTracer{} return &nopTracer{}
@ -70,3 +149,12 @@ type nopSpan struct{}
func (s *nopSpan) Finish() {} func (s *nopSpan) Finish() {}
func (s *nopSpan) LogKV(alternatingKeyValues ...interface{}) {} func (s *nopSpan) LogKV(alternatingKeyValues ...interface{}) {}
type arbitraryContextKeyType int
var arbitraryContextKey arbitraryContextKeyType
func spanFromContext(ctx context.Context) (Span, bool) {
span, ok := ctx.Value(arbitraryContextKey).(Span)
return span, ok
}

View file

@ -301,6 +301,9 @@ func (t *ClusterCluster) Close() error {
if err != nil { if err != nil {
return err return err
} }
// Make sure open indexes get shut down too. we wouldn't do
// this normally for a cluster, but we want to for test cases.
c.holder.Close()
} }
return nil return nil
} }

18
view.go
View file

@ -59,7 +59,7 @@ type view struct {
stats stats.StatsClient stats stats.StatsClient
rowAttrStore AttrStore rowAttrStore AttrStore
logger logger.Logger logger logger.Logger
snapshotQueue chan *fragment snapshotQueue snapshotQueue
} }
// newView returns a new instance of View. // newView returns a new instance of View.
@ -210,7 +210,7 @@ fragLoop:
// flags returns a set of flags for the underlying fragments. // flags returns a set of flags for the underlying fragments.
func (v *view) flags() byte { func (v *view) flags() byte {
var flag byte var flag byte
if v.fieldType == FieldTypeInt { if v.fieldType == FieldTypeInt || v.fieldType == FieldTypeDecimal {
flag |= roaringFlagBSIv2 flag |= roaringFlagBSIv2
} }
return flag return flag
@ -309,7 +309,9 @@ func (v *view) newFragment(path string, shard uint64) *fragment {
frag.CacheSize = v.cacheSize frag.CacheSize = v.cacheSize
frag.Logger = v.logger frag.Logger = v.logger
frag.stats = v.stats frag.stats = v.stats
if v.snapshotQueue != nil {
frag.snapshotQueue = v.snapshotQueue frag.snapshotQueue = v.snapshotQueue
}
if v.fieldType == FieldTypeMutex { if v.fieldType == FieldTypeMutex {
frag.mutexVector = newRowsVector(frag) frag.mutexVector = newRowsVector(frag)
} else if v.fieldType == FieldTypeBool { } else if v.fieldType == FieldTypeBool {
@ -403,6 +405,16 @@ func (v *view) setValue(columnID uint64, bitDepth uint, value int64) (changed bo
return frag.setValue(columnID, bitDepth, value) return frag.setValue(columnID, bitDepth, value)
} }
// clearValue removes a specific value assigned to columnID
func (v *view) clearValue(columnID uint64, bitDepth uint, value int64) (changed bool, err error) {
shard := columnID / ShardWidth
frag := v.Fragment(shard)
if frag == nil {
return false, nil
}
return frag.clearValue(columnID, bitDepth, value)
}
// sum returns the sum & count of a field. // sum returns the sum & count of a field.
func (v *view) sum(filter *Row, bitDepth uint) (sum int64, count uint64, err error) { func (v *view) sum(filter *Row, bitDepth uint) (sum int64, count uint64, err error) {
for _, f := range v.allFragments() { for _, f := range v.allFragments() {
@ -483,7 +495,7 @@ func upgradeViewBSIv2(v *view, bitDepth uint) (ok bool, _ error) {
if tmpPath, err := upgradeRoaringBSIv2(frag, bitDepth); err != nil { if tmpPath, err := upgradeRoaringBSIv2(frag, bitDepth); err != nil {
return ok, errors.Wrap(err, "upgrading bsi v2") return ok, errors.Wrap(err, "upgrading bsi v2")
} else if err := frag.closeStorage(true); err != nil { } else if err := frag.closeStorage(); err != nil {
return ok, errors.Wrap(err, "closing after bsi v2 upgrade") return ok, errors.Wrap(err, "closing after bsi v2 upgrade")
} else if err := os.Rename(tmpPath, frag.path); err != nil { } else if err := os.Rename(tmpPath, frag.path); err != nil {
return ok, errors.Wrap(err, "renaming after bsi v2 upgrade") return ok, errors.Wrap(err, "renaming after bsi v2 upgrade")