The IDK tests should be run only when there's changes in the IDK
or client directories.
Also the shard transactional test should be run ever.
The "optional: true" flag is a fascinating quirk allowing you to
express that, *if* a job exists, we should wait for it, but if the
job doesn't exist, that's also acceptable.
When we make dummy test servers, we should make them using sockets
for etcd rather than TCP ports so we don't run into problems like
the test always failing if anything else is on that port already,
which it can totally legitimately be. For instance, if you ran
an existing featurebase server, and then tried "go test" in the
server directory, this would fail.
We want "make test" to run in a reasonable amount of time and
actually work, the IDK tests are full of tests that only run in
a specialized docker environment with things like hosts named
pilosa and kafka and such.
* FB-1251: Add ability to sort Extract queries by some field
There is a sort call which takes a row call and the field, and
based on the field type, the corresponding rows are read. both key and
value are stored in RowKV{}. The value is stored since its required to
merge data from shards. values are sorted in each shard and these sorted
listes are merged in the reduce.
sort-desc flag is sent to comparator to decide the sorting order. ok
flag is added to the compare function to track any error in the
sort.Slice anonymous function
Sorting over set field was removed, since there would be multiple values
for each ids and there would be no right sorting order there.
Before, we were building docker images for IDK for each of the four
linux/darwin amd64/arm64 platfrom/arch combinations, which didn't make
sense. If we want to later build docker images for linux/arm64, we can
add that later.
I also cleaned up the Dockerfile for IDK to minimize creation of excess
layers (by &&-ing RUN commands), and made apt quieter to cut back some
of the noise.
We need to install base gosec tool and then the gitlab version to convert the gosec
json to the gl-sast-report.json that GitLab expects.
This lets us see the 'Security' tab under pipelines (and under the default branch
after this change is merged).
I chose to pin both of the versions of the tools to avoid any dependencies changing.
This could be an issue, but both repos are largely frozen.
We're using this because it's a builtin env variable that comes
with GitLab, and it fixes one of the annoying things about Docker
tags (e.g., you can't use all of the allowed characters in
Git branches).
One issue that I've seen a few times, is branches with either
capital letters (which was recently broken), or using the '/'
character.
This PR makes it so we always use the CI_COMMIT_REF_SLUG
when making or referencing images so that it is always consistent.
Note: this might make it slightly harder to intuit what the correct
Docker image to make (if you wanted to use the one built by CI rather
than locally). This trade off doesn't seem too hard to overcome.
There are currently three copies of a package called `fakeidp` in the
featurebase repo:
- ./idk/fakeidp/go.mod
- ./internal/clustertests/fakeidp/go.mod
- ./qa/fakeidp/go.mod
All three have a `go.mod` file. While this is supported under golang's
new Workspace support, what's not supported is that the modules share
the same name (in this case "fakeidp"). This commit is a sort of
temporary fix which renames the module for two of the instances. This
prevents, for example, VSCode with workspace support enabled, from
barfing.
By the way, one can enable VSCode workspace support with the following
setting:
```
// gopls
"gopls": {
"build.experimentalWorkspaceModule": true
}
```
Also...
This commit fixes the `make testv` target. It's probably not used
anywhere (which I'm assuming because it was broken), but it's a handle
target, so now it will list and run tests against all packages found in
the repo, including the root package.
* [CLOUD-934] Optionally broadcast IDK Kinesis errors/panics to external storage
- Add a minor public method `idk.Main.SetLog` to allow setting the logger instance
after initialization.
- Add a Logger implementation that captures recoverable errors and panics
and pushes to an external store. Meant to decorate an existing Logger
instance and always delegate to its implementation. Decoration happens
when all AWS resources are initialized. Before then, the wrapped Logger
implementation is used.
- If `--error-queue-name/CONSUMER_ERROR_QUEUE_NAME` specified, use an
ErrorStreamLogger to push errors and panics to an SQS queue with that name.
Omission of the option preserves current behavior.
- Parse sink ID from the `--stream-name/CONSUMER_STREAM_NAME` expecting the form
'PREFIX'-VALID_UUID. If the sink UUID is invalid, emit a warning that errors/panics
will not be written to an SQS queue but will still be logged using the decorated
Logger instance.
- The inability to push to an SQS queue leads to warnings being emitted to notify
ECS that no queue will be written to and is NOT a hard error.
- Add SQS interface mock for unit testing.
- Add IDK make targets for generating mock interfaces.
* [CLOUD-934] Execute go mod tidy and go fmt to pass CI/CD checks
* [CLOUD-934] Remove extraneous Makefile in idk/kinesis and fix install-mock-generator target
* [CLOUD-934] Add godocs to exported types and functions
* [CLOUD-934] Changed warning to not sound so ominous and update associated unit test
* [CLOUD-934] Unblock CI/CD at the IDK test stage
previously we allowed users to specify a granularity for timestamp
e.g. seconds, milli, micro, nano
however we converted everything to nano before we stored it.
This reduced the allowed range for all time units to what
was allowed by timestamp. For example, with second granularity
you can represent billions of years within the capacity of
int64 but with nano its somewhere b/w 100-200 years.
So now, for timeunits of seconds, milli, and micro the range
is year 0001 - 9999. These limits come from what Go
supports.
So this uses unit specific function to translate
timestamps to values and vice versa to increase
the time range.
In the process of increasing the range for timestamp and subsequent
testing, I found and addressed a few bugs:
- min/max queries were not using timestamp specific comparators so
added that.
- Values from Import/ingest come to FB as relative values to epoch
whereas other BSI fields come as actual values and then
becomes relative to their respective bases within FB. so some
specific handling of that was added.
- However! Set queries use timestamp strings which are, of course,
the actual value they designate. So they have to become
relative.
- When bitdepth is 0, Min/maxUnsigned functions did not run
resulting in a count of 0 when there
was an actual value that was 0.
Also, this removes (now) dead code and updates/adds tests.
most types are imported in the format `<row>,<col>`, but ints and decimals
aren't. with this new flag, ints and decimals are imported using the
`<row>,<col>` format, instead of `<col>,<row>`.
compares free space in output directory to
the usage of either the data directory or
index depending on what is being backed up.
- adds an http_handler endpoint to get usage
of a particular index
- adds InternalClient methods to get DiskUsage and
IndexUsage
* unifying idk and featurebase: first pass
* resolved conflict with master for gitignore & dockerignore
* deleted binaries that were accidentally pushed to git
* combined gitlab jobs for idk & featurebase
* run go fmt for idk
* updated ssh env variable, and made docker password variable in gitlab env variables
* fixed typo assigning variable name
* trying to fix docker login error
* trying a different solution for docker password
* pass registry
* fixed docker login
* updated paths for idk
* exclude idk tests from featurebase test run
* fix vendor error
* update certificates
* grpc needs to be in version 1.38
genproto, which is imported by big query updates the grpc version to 1.47.0
grpc 1.47.0 causes etcd to deadlock when calling etcd.Close()
the fix is to have a replace in go.mod to specify a specific grpc version
* run go mod tidy
* go mod
* run go mod tidy
* exclude bigquery since it is causing issues and undo grpc replace in go.mod
* fix grpc version
* fix formatting error
* update formatting
* attempt to fix formatting
* update path for code coverage
* update to use current branch binaries, not master
* fix for building idk - path updates
* udpate path for binaries
* update job dependecies
* update docker idk tests to use the current branch registry
* update stages for jobs
* updated job dependencies
* not allow idk s3 dump to fail since it is a dependency for integration tests
* update dependecy for idk tests
* update paths for idk build and code coverage
* download featurebase binary from s3
* pass branch name to all setup scripts
* change to current branch instead of master
* updated sonarcloud
* sonarcloud fix and branch name fix
* trying to speed up pipeline run time
* update stage
* branch name fix + sonar cloud
* sonarcloud
* Fixed broken fields when packaging rpm and deb files.
* Update systemd unit files, package them into RPM's.
* Fix config file path for packages.
* Changed unit and binary paths to conform to standard locations for each vendor.
* Added featurebase owned directories.
* Create featrebase user/group and chown the right dirs
* Automate turning on featurebase
* Updated the .gitignore to include .vscode files.
* Changed RPM name to better conform to naming standards.
* Pass GOARCH when building RPM's.
* Avoid using recursive to remove files in this dir.
In the logs, I can see that this error occurs when a query is done
during the delete view.This fix is only to bypass it and log that
data.
The real issue is that the delete standard view which should happen
only once, is occuring every hour or two. The ingester might be
creating the standard views which needs to be fixed.
The root problem this is attempting to address is sporadic
weird cases in which etcd mistakenly thinks it's down even when
it's up. I am not confident that this is addressed, but there's
a reasonable chance that it is, and I can't trigger it at the
moment, but it was always sporadic, so that doesn't prove much.
There's a lot going on here, and it comes into roughly three
categories.
First: Dropping unused/unneeded code. There's a lot of leftover
bits from the initial development and refactoring of this.
Second: Unifying and shuffling some of the design. We had
multiple interfaces which are functionally impossible to
usefully implement separately, so they're combined together,
and in some cases, moved.
Third: Streamlining logic and simplifying design choices.
This is combined into one commit because the changes are
thoroughly entertwined with each other and you can't usefully
break most of them out.
Also, a bunch of test coverage for most of these changes.
Big changes:
We merge the topology and disco packages. The topology and disco
packages being separate creates a complicated tangle of problems
and dependencies. The fundamental problem, approximately, is that
topology.Node has to track disco.NodeState.
There's three core interfaces interacting here:
topology.Noder (maintains list of nodes)
disco.Stator (maintains the state of a node)
disco.Metadator (stores, possibly retrieves, node metadata)
But the node state mantained by the Noder *is* the set of node
metadata, plus state updates produced by Stators. The only actual
non-trivial and usable implementation of these interfaces is a single
thing which implements all three, and in which the implementations
share a single backend data source which they are all modifying.
But you can't move Noder into disco, because Noder has to refer
to topology.Node, but topology.Node refers to disco.
Solution: First, merge these two packages. Second, merge these
three interfaces, to provide a single interface which is more
clear about the fact that (metadator.)SetMetadata() and
(stator.)Started() are both changing the output we'll get from
(noder.)Nodes().
We rework the node state tracking.
We have this nodeStates map which is almost unused. Really, we
don't need it at all. Every node's state is either its last heartbeat
state or "Unknown", so we simplify this a bit. Also, we ensure that
the populateNodeStates function itself is yielding the sorted nodes
list, so we don't have to be as worried about possible later lookups
of sortedNodes happening outside a lock. We also add diagnostics
for deleting nodes from the metadata list (this should never happen),
and try to track heartbeat state more closely.
This is *probably* what fixes the underlying reported problem,
if anything did.
Still an open issue: Make heartbeat state changes aware of when
they're talking about *this* node and possibly not try to
mark it down? Except this may have a flaw: That would result in
each node disagreeing with other nodes in etcd about the state
of that node in the failure cases, and undermine the point of
using etcd to keep these states consistent.
We reduce the number of contexts and cancelfuncs in the etcd wrapper.
We create a shared context for the non-etcd.embed children of our
etcd wrapper, the heartbeat/keepalive and the node watcher, so we
can cancel that one context and cancel all of those at once, so
we don't need to separately track a function to call to cancel
the watch, AND be closing another channel. Also, our shutdown
now propagates automatically to the various etcd API calls we've
made for things like the node watcher and keepalive calls.
We still need to watch that channel in watchNodesOnce, though,
because apparently the watch doesn't yield an error even if the
context calling it is canceled. Whee.
This should reduce the risk of ending up in an inconsistent state,
and also the Close() function is probably idempotent now.
Smaller changes:
* Remove config-generators that existed to generate etcd
configs but were used only for tests that no longer exist
or make sense.
* Move the logic to generate etcd configs into the etcd
package, instead of the "testing" subpackage. This allows
us to write a self-contained config generator for
clusters where the nodes know about each other, but do
this just with etcd, not with full featurebase servers.
* Move the thing generating `fake:%d` socket names into
the etcd package, which is the only place we use it.
Also simplify it slightly.
* Don't panic on invalid URLs, report errors from them.
* At least try to use etcd's config.Validate functionality.
It's underdocumented, so we're not sure what it will report,
but at least if it does we'll get reports from it and
know what they are?
* Try to handle CompactRevision errors from watches more
correctly -- after a CompactRevision, any future attempt
to watch from a lower revision will necessarily fail, so
we adjust our target revision up. We don't have good
testing for this.
* Drop the Metadata() method (that used to be in Metadator)
because nothing ever used it and it didn't make much sense
to try.
* Convert SetMetadata from taking an arbitrary json blob
to taking the only data that would ever be valid since
we always use it to extract node information anyway.
* Drop several unused functions, unexport things only used
internally.
* Replace Started() with SetState("STARTED"), allowing us
to write tests that mess with states. We weren't really thinking
carefully about state transitions sometimes and now it's much
easier to do that thinking.
* Stop leaving stray localhost:2380 and localhost:2379 in
our embed config. We still sometimes see peer requests from
those and I honestly don't know why, but at least it should
be rarer.
Changes to row queries with from/to options return an error if field is not a timestamp field
logic couldn't be made earlier in call stack because other queries process from/to time differently.
This PR reverts the timestamp work.
The timestamp work requires changes to FB and IDK; and there are
circular dependencies between tests in either repo preventing merging
of either. The work here is pretty stable, but required bypassing
the smoketest. Meanwhile, I found some additional things in IDK that
need addressing which means I merged this work in pre-maturely. Once
I get that worked out, I'll re-commit these commits.
This reverts following commits related to timestamp work:
bypass of smoke test b/c of circular dep with IDK: 0676790
update codec to reflect changes to timstamp range: bd5dc76
fix few bugs regarding timestamp: 33fce8a
increase time range for timestamp by using specified granularity: 5939923.
In the process of increasing the range for timestamp and subsequent
testing, I found and addressed a few bugs:
- min/max queries were not using timestamp specific comparators so
added that.
- Values from Import/ingest come to FB as relative values to epoch
whereas other BSI fields come as actual values and then
becomes relative to their respective bases within FB. so some
specific handling of that was added.
- However! Set queries use timestamp strings which are, of course,
the actual value they designate. So they have to become
relative.
- When bitdepth is 0, Min/maxUnsigned functions did not run
resulting in a count of 0 when there
was an actual value that was 0.
Also, this removes (now) dead code and updates/adds tests.
previously we allowed users to specify a granularity for timestamp
e.g. seconds, milli, micro, nano
however we converted everything to nano before we stored it.
This reduced the allowed range for all time units to what
was allowed by timestamp. For example, with second granularity
you can represent billions of years within the capacity of
int64 but with nano its somewhere b/w 100-200 years.
So now, for timeunits of seconds, milli, and micro the range
is year 0001 - 9999. These limits come from what Go
supports.
So this uses unit specific function to translate
timestamps to values and vice versa to increase
the time range.